SYSTEM ONE MODEL INTELLIGENCE DIRECTORY|Community workflows and reports
Independent site
Last updated:

Jev benchmark reports, with their limits

These numbers come from the cited authors and repositories. We checked the sources but did not rerun their API workloads. Different datasets and pipelines appear in separate tables so they are not mistaken for a model leaderboard.

Business-email classification

Nikhil Mudholkar: 1,565 emails across 10 categories, including 1,201 real emails and 364 AI-written examples.

Results reported by the author: Business-email classification
ModelOverall accuracyEqual-weight category accuracyCost per 1,000 emails
Jev96.4%92.0%$0.08
Gemini 3.5 Flash-Lite97.5%94.6%$0.80
Gemini 3.8 Flash98.5%96.9%$1.79
  • These are author-reported results against the reference labels. The comparison runs each model separately; it does not evaluate a combined pipeline.
  • The emails are mostly German, with English examples. Attachments are excluded and the AI-written examples target rare categories.
  • Cost is normalized to 1,000 emails, not the total cost of the dataset. The author notes that the savings are small at normal inbox volumes.

Sources:[9] nikhil mudholkar[10] nikhil mudholkar[11] nikhil mudholkar[13] nikhil mudholkar[14] nikhil mudholkar

Read the model report and user feedback

Labeling 1,000 app reviews

Rahul Kumar’s Column Race: sentiment, topic, bug and churn labels for 1,000 Android reviews. Both lanes use batches of 20 at concurrency 8.

Results reported by the author: Labeling 1,000 app reviews
ModelWhole-run durationEstimated run costSentiment correlation with stars
Jev 1.13.04.6 seconds$0.0230.803
Gemini 3.8 Flash18.8 seconds$0.1580.822
  • The public repository contains one recorded run pair with these results. Durations cover all 1,000 reviews and are not per-request medians.
  • Sentiment uses Spearman correlation with star ratings, a weak proxy. Topic agreement of 84.4% and bug agreement of 92.6% have no human ground truth and must not be labeled accuracy.
  • The rounded costs use reported token counts and list prices, not invoices. No JSON retries occurred in this recorded pair.

Sources:[15] Rahul Kumar[16] goodrahstar

Read the model report and user feedback

Email screening with a Kimi K3 fallback

Hassan: 100 emails, evenly split between legitimate and fraudulent examples. Jev sends 31 cases below 95% confidence to Kimi K3.

Results reported by the author: Email screening with a Kimi K3 fallback
WorkflowWhole-run durationReported run costCorrect decisions
Jev + Kimi K316 secondsAbout $0.0796 of 100
  • Jev’s first pass takes 1.42 seconds. That excludes Kimi’s work and must not be presented as the full pipeline duration.
  • The author attributes $0.003 to Jev and $0.068 to Kimi. The roughly $0.07 total is rounded.
  • This is an author-reported demonstration, not a production fraud evaluation or a comparison against the two Gemini datasets. The 95% confidence cutoff is specific to the demo.

Sources:[19] Hassan

Read the model report and user feedback

Sources and methodology

nikhil mudholkar (@nikhilmudholkar)

Published:

Source checked:

X.com

[9] Business-email classification against Gemini

Mudholkar compares Jev with Gemini on 1,565 German and English business emails in 10 categories. Gemini performs better on overall accuracy; the author is interested in Jev for its uncertainty signal.

An author-reported benchmark. It compares separate models, rather than measuring a combined Jev + Gemini pipeline.

Jev / GeminiRead original

nikhil mudholkar (@nikhilmudholkar)

Published:

Source checked:

X.com

[10] Email dataset composition and shared instructions

The dataset contains 1,201 real emails and 364 AI-written examples for rare categories, mostly in German. The three models receive the same email text and category instructions.

Methodology from the benchmark thread. The dataset is not entirely real correspondence.

Jev / GeminiRead original

nikhil mudholkar (@nikhilmudholkar)

Published:

Source checked:

X.com

[11] Email accuracy results for Jev and two Gemini models

Reported overall accuracy is 96.4% for Jev, 97.5% for Gemini 3.5 Flash-Lite and 98.5% for Gemini 3.8 Flash. Equal weighting of categories yields 92.0%, 94.6% and 96.9%, respectively.

Results against the author’s reference labels in this dataset. These are not general accuracy scores for the models.

Jev / GeminiRead original

nikhil mudholkar (@nikhilmudholkar)

Published:

Source checked:

X.com

[13] Reported cost per 1,000 business emails

The author reports $0.08 for Jev, $0.80 for Gemini 3.5 Flash-Lite and $1.79 for Gemini 3.8 Flash per 1,000 emails in these runs.

Workload-specific costs, not provider pricing or the cost of the whole 1,565-email dataset. The author notes that savings are small at ordinary inbox volumes.

Jev / GeminiRead original

nikhil mudholkar (@nikhilmudholkar)

Published:

Source checked:

X.com

[14] Attachments and synthetic examples in the email test

Attachments were excluded, and 364 emails were generated to match chosen labels. The author describes the lack of non-text input as a barrier to production use.

Limitations supplied by the experiment’s author. This source does not demonstrate processing email attachments.

Jev / GeminiRead original

Rahul Kumar (@rahulbuildsmore)

Published:

Source checked:

X.com

[15] Jev and Gemini labeling 1,000 app reviews

Kumar compares sentiment, topic, bug and churn labels on 1,000 app reviews. He reports Jev at 4.6 seconds and $0.023, versus Gemini 3.8 Flash at 18.8 seconds and $0.158.

The post links to a public project with recorded runs and measurement notes. It is a side-by-side comparison, not a combined workflow.

Jev / GeminiRead original

goodrahstar

Source checked:

GitHub

[16] Column Race recorded runs and evaluation methodology

Both lanes process 20 reviews per request at concurrency 8. The recorded pair uses Jev 1.13.0 and Gemini 3.8 Flash. Sentiment correlation with star ratings is 0.803 and 0.822; topic agreement is 84.4% and bug agreement is 92.6%.

One recorded pair over 1,000 Android reviews. Stars are a weak sentiment proxy; topic and bug labels have no human ground truth. Agreement is not accuracy. Costs use token counts and list prices, not invoices.

Jev / GeminiRead original

Hassan (@nutlope)

Published:

Source checked:

X.com

[19] Jev first-pass email screening with Kimi K3 fallback

Hassan tests 50 legitimate and 50 fraudulent emails. Jev classifies them in 1.42 seconds; 31 cases below a 95% confidence threshold go to Kimi K3. The full pipeline takes 16 seconds, gets 96 of 100 correct and costs about $0.07.

Author-reported demonstration, not independently rerun here. The stated cost split is $0.003 for Jev and $0.068 for Kimi. The 1.42-second figure excludes the Kimi stage.

Jev / KimiRead original