Jev benchmark reports, with their limits
These numbers come from the cited authors and repositories. We checked the sources but did not rerun their API workloads. Different datasets and pipelines appear in separate tables so they are not mistaken for a model leaderboard.
Business-email classification
Nikhil Mudholkar: 1,565 emails across 10 categories, including 1,201 real emails and 364 AI-written examples.
| Model | Overall accuracy | Equal-weight category accuracy | Cost per 1,000 emails |
|---|---|---|---|
| Jev | 96.4% | 92.0% | $0.08 |
| Gemini 3.5 Flash-Lite | 97.5% | 94.6% | $0.80 |
| Gemini 3.8 Flash | 98.5% | 96.9% | $1.79 |
- These are author-reported results against the reference labels. The comparison runs each model separately; it does not evaluate a combined pipeline.
- The emails are mostly German, with English examples. Attachments are excluded and the AI-written examples target rare categories.
- Cost is normalized to 1,000 emails, not the total cost of the dataset. The author notes that the savings are small at normal inbox volumes.
Sources:[9] nikhil mudholkar[10] nikhil mudholkar[11] nikhil mudholkar[13] nikhil mudholkar[14] nikhil mudholkar
Read the model report and user feedbackLabeling 1,000 app reviews
Rahul Kumar’s Column Race: sentiment, topic, bug and churn labels for 1,000 Android reviews. Both lanes use batches of 20 at concurrency 8.
| Model | Whole-run duration | Estimated run cost | Sentiment correlation with stars |
|---|---|---|---|
| Jev 1.13.0 | 4.6 seconds | $0.023 | 0.803 |
| Gemini 3.8 Flash | 18.8 seconds | $0.158 | 0.822 |
- The public repository contains one recorded run pair with these results. Durations cover all 1,000 reviews and are not per-request medians.
- Sentiment uses Spearman correlation with star ratings, a weak proxy. Topic agreement of 84.4% and bug agreement of 92.6% have no human ground truth and must not be labeled accuracy.
- The rounded costs use reported token counts and list prices, not invoices. No JSON retries occurred in this recorded pair.
Sources:[15] Rahul Kumar[16] goodrahstar
Read the model report and user feedbackEmail screening with a Kimi K3 fallback
Hassan: 100 emails, evenly split between legitimate and fraudulent examples. Jev sends 31 cases below 95% confidence to Kimi K3.
| Workflow | Whole-run duration | Reported run cost | Correct decisions |
|---|---|---|---|
| Jev + Kimi K3 | 16 seconds | About $0.07 | 96 of 100 |
- Jev’s first pass takes 1.42 seconds. That excludes Kimi’s work and must not be presented as the full pipeline duration.
- The author attributes $0.003 to Jev and $0.068 to Kimi. The roughly $0.07 total is rounded.
- This is an author-reported demonstration, not a production fraud evaluation or a comparison against the two Gemini datasets. The 95% confidence cutoff is specific to the demo.
Sources:[19] Hassan
Read the model report and user feedbackSources and methodology
nikhil mudholkar (@nikhilmudholkar)
Published:
Source checked:
[9] Business-email classification against Gemini
Mudholkar compares Jev with Gemini on 1,565 German and English business emails in 10 categories. Gemini performs better on overall accuracy; the author is interested in Jev for its uncertainty signal.
An author-reported benchmark. It compares separate models, rather than measuring a combined Jev + Gemini pipeline.
nikhil mudholkar (@nikhilmudholkar)
Published:
Source checked:
[10] Email dataset composition and shared instructions
The dataset contains 1,201 real emails and 364 AI-written examples for rare categories, mostly in German. The three models receive the same email text and category instructions.
Methodology from the benchmark thread. The dataset is not entirely real correspondence.
nikhil mudholkar (@nikhilmudholkar)
Published:
Source checked:
[11] Email accuracy results for Jev and two Gemini models
Reported overall accuracy is 96.4% for Jev, 97.5% for Gemini 3.5 Flash-Lite and 98.5% for Gemini 3.8 Flash. Equal weighting of categories yields 92.0%, 94.6% and 96.9%, respectively.
Results against the author’s reference labels in this dataset. These are not general accuracy scores for the models.
nikhil mudholkar (@nikhilmudholkar)
Published:
Source checked:
[13] Reported cost per 1,000 business emails
The author reports $0.08 for Jev, $0.80 for Gemini 3.5 Flash-Lite and $1.79 for Gemini 3.8 Flash per 1,000 emails in these runs.
Workload-specific costs, not provider pricing or the cost of the whole 1,565-email dataset. The author notes that savings are small at ordinary inbox volumes.
nikhil mudholkar (@nikhilmudholkar)
Published:
Source checked:
[14] Attachments and synthetic examples in the email test
Attachments were excluded, and 364 emails were generated to match chosen labels. The author describes the lack of non-text input as a barrier to production use.
Limitations supplied by the experiment’s author. This source does not demonstrate processing email attachments.
Rahul Kumar (@rahulbuildsmore)
Published:
Source checked:
[15] Jev and Gemini labeling 1,000 app reviews
Kumar compares sentiment, topic, bug and churn labels on 1,000 app reviews. He reports Jev at 4.6 seconds and $0.023, versus Gemini 3.8 Flash at 18.8 seconds and $0.158.
The post links to a public project with recorded runs and measurement notes. It is a side-by-side comparison, not a combined workflow.
goodrahstar
Source checked:
[16] Column Race recorded runs and evaluation methodology
Both lanes process 20 reviews per request at concurrency 8. The recorded pair uses Jev 1.13.0 and Gemini 3.8 Flash. Sentiment correlation with star ratings is 0.803 and 0.822; topic agreement is 84.4% and bug agreement is 92.6%.
One recorded pair over 1,000 Android reviews. Stars are a weak sentiment proxy; topic and bug labels have no human ground truth. Agreement is not accuracy. Costs use token counts and list prices, not invoices.
Hassan (@nutlope)
Published:
Source checked:
[19] Jev first-pass email screening with Kimi K3 fallback
Hassan tests 50 legitimate and 50 fraudulent emails. Jev classifies them in 1.42 seconds; 31 cases below a 95% confidence threshold go to Kimi K3. The full pipeline takes 16 seconds, gets 96 of 100 correct and costs about $0.07.
Author-reported demonstration, not independently rerun here. The stated cost split is $0.003 for Jev and $0.068 for Kimi. The 1.42-second figure excludes the Kimi stage.