Sources behind the Jev reports
We read the original X posts, follow-up discussions and linked project documentation before publishing. These cards are editorial summaries, not quotations. A source check confirms what the author reported; it does not mean we reproduced the experiment. Measurements retain their task, model and sample limits.
Source notes (24)
TypeSafe AI
Source checked:
[1] Jev and the Choice, Score and Noul primitives
The documentation describes Jev as a System One model that evaluates state and typed questions. Choice selects an option, Score evaluates a rubric, and Noul returns a probability for a statement.
API behavior documented by the provider. This page does not establish a universal latency, accuracy or cost advantage.
TypeSafe AI
Source checked:
[2] Confidence, probabilities and fallback decisions
Choice and Score include confidence derived from their probability distributions. Noul does not have a separate confidence field. TypeSafe recommends choosing thresholds for the domain and consequences of an action.
Implementation reference. A confidence value is not a guarantee that an individual decision is correct.
paulwei (@coolish)
Published:
Source checked:
[3] Trying Jev after GPT-6 Astra in Slay the Spire 2
The author describes GPT-6 Astra as capable but slow in an earlier game run, then reports roughly 0.7 seconds of action deliberation with Jev.
Personal game experiment with a video. The timing is not a controlled, repeated GPT-versus-Jev benchmark.
paulwei (@coolish)
Published:
Source checked:
[4] Astra revises Jev prompts after a failed game run
The first Ascension 10 attempt reached floor 6. The author says Astra then iterated on Jev prompts and the pair reached floor 17, while describing Jev as weaker at the game than Astra.
A follow-up from the same experiment, not independent corroboration. Reaching floor 17 is not a reported completed run.
Sac (@Saccc_c)
Published:
Source checked:
[5] Codex and Jev for adding a Mac calendar event
Sac describes a Codex + Jev computer-use experiment and compares adding the same Mac calendar event. The reported experience is smoother, with similar token consumption.
Qualitative report and comparison video. The post does not identify the underlying GPT version or give a repeatable latency measurement.
tamara (@tamarajtran)
Published:
Source checked:
[6] Scoring tool calls for context compaction
The project author proposes using Jev to score tool calls and discard irrelevant material during compaction.
Project announcement. The repository says its native demo app is scripted and makes no API calls, so its animation is not timing evidence.
Alex Volkov (@altryne)
Published:
Source checked:
[7] A user report of Claude context reduction
Volkov reports reducing a Claude session from nearly one million tokens to about 86,000 in roughly one second using fast-jev-compaction.
Individual experience, not a controlled benchmark or a measure of downstream task accuracy. The timestamp is the edited-post time shown by X.
tamaratran
Source checked:
[8] fast-jev-compaction implementation and limitations
The plugin scores tool calls and results, preserves user and assistant text in its output, and keeps calls paired with their results. Its hook falls back to built-in compaction on errors or insufficient reduction.
The README documents early-access Claude Code function hooks, estimated token sizes and a scripted native demo. It explicitly warns that a probability does not prove a result is safe to delete.
nikhil mudholkar (@nikhilmudholkar)
Published:
Source checked:
[9] Business-email classification against Gemini
Mudholkar compares Jev with Gemini on 1,565 German and English business emails in 10 categories. Gemini performs better on overall accuracy; the author is interested in Jev for its uncertainty signal.
An author-reported benchmark. It compares separate models, rather than measuring a combined Jev + Gemini pipeline.
nikhil mudholkar (@nikhilmudholkar)
Published:
Source checked:
[10] Email dataset composition and shared instructions
The dataset contains 1,201 real emails and 364 AI-written examples for rare categories, mostly in German. The three models receive the same email text and category instructions.
Methodology from the benchmark thread. The dataset is not entirely real correspondence.
nikhil mudholkar (@nikhilmudholkar)
Published:
Source checked:
[11] Email accuracy results for Jev and two Gemini models
Reported overall accuracy is 96.4% for Jev, 97.5% for Gemini 3.5 Flash-Lite and 98.5% for Gemini 3.8 Flash. Equal weighting of categories yields 92.0%, 94.6% and 96.9%, respectively.
Results against the author’s reference labels in this dataset. These are not general accuracy scores for the models.
nikhil mudholkar (@nikhilmudholkar)
Published:
Source checked:
[12] Where the email-classification mistakes occurred
The author reports that 737 predictions at 99% confidence or above matched the reference labels. Below 70% confidence, nearly half were wrong.
An observed relationship in one dataset, not a universal threshold or proof of error-free decisions.
nikhil mudholkar (@nikhilmudholkar)
Published:
Source checked:
[13] Reported cost per 1,000 business emails
The author reports $0.08 for Jev, $0.80 for Gemini 3.5 Flash-Lite and $1.79 for Gemini 3.8 Flash per 1,000 emails in these runs.
Workload-specific costs, not provider pricing or the cost of the whole 1,565-email dataset. The author notes that savings are small at ordinary inbox volumes.
nikhil mudholkar (@nikhilmudholkar)
Published:
Source checked:
[14] Attachments and synthetic examples in the email test
Attachments were excluded, and 364 emails were generated to match chosen labels. The author describes the lack of non-text input as a barrier to production use.
Limitations supplied by the experiment’s author. This source does not demonstrate processing email attachments.
Rahul Kumar (@rahulbuildsmore)
Published:
Source checked:
[15] Jev and Gemini labeling 1,000 app reviews
Kumar compares sentiment, topic, bug and churn labels on 1,000 app reviews. He reports Jev at 4.6 seconds and $0.023, versus Gemini 3.8 Flash at 18.8 seconds and $0.158.
The post links to a public project with recorded runs and measurement notes. It is a side-by-side comparison, not a combined workflow.
goodrahstar
Source checked:
[16] Column Race recorded runs and evaluation methodology
Both lanes process 20 reviews per request at concurrency 8. The recorded pair uses Jev 1.13.0 and Gemini 3.8 Flash. Sentiment correlation with star ratings is 0.803 and 0.822; topic agreement is 84.4% and bug agreement is 92.6%.
One recorded pair over 1,000 Android reviews. Stars are a weak sentiment proxy; topic and bug labels have no human ground truth. Agreement is not accuracy. Costs use token counts and list prices, not invoices.
Dan McAteer (@daniel_mac8)
Published:
Source checked:
[17] Ranking X posts with Jev and Grok Bot
McAteer describes retrieving 1,000 agent-related posts with the X API, using Jev for pairwise selection, and having Grok Bot surface five tips.
A first-hand workflow description. No model version, comparison count, latency or cost breakdown is supplied.
codila (@0xCodila)
Published:
Source checked:
[18] A Jev router in front of GrokBot actions
The author describes consulting Jev before browser, research and retry actions, starting in shadow mode, inspecting logs, and retaining a bypass and human control over irreversible actions.
A personal setup report with promotional language. Its broad performance claims and seven-minute setup claim are not treated as measurements.
Hassan (@nutlope)
Published:
Source checked:
[19] Jev first-pass email screening with Kimi K3 fallback
Hassan tests 50 legitimate and 50 fraudulent emails. Jev classifies them in 1.42 seconds; 31 cases below a 95% confidence threshold go to Kimi K3. The full pipeline takes 16 seconds, gets 96 of 100 correct and costs about $0.07.
Author-reported demonstration, not independently rerun here. The stated cost split is $0.003 for Jev and $0.068 for Kimi. The 1.42-second figure excludes the Kimi stage.
yonsakhan (@yonsakhan)
Published:
Source checked:
[20] Using Kimi K3 to develop a Jev-powered comment filter
The author says they developed an X comment-filtering userscript with Kimi K3 and use Jev for semantic spam decisions, including obfuscated text. They describe reversible collapsing and a local cache.
A development report. Kimi helped write the software; the post does not establish that Kimi participates in its runtime classification loop.
Jing Wang (@jingwangtalk)
Published:
Source checked:
[21] A model-routing experiment sensitive to the candidate list
Wang experiments with a Jev router over DeepSeek, GLM, Kimi, GPT and Claude candidates. Adding Claude Opus 5 changes the selected model from DeepSeek Flash to Claude Sonnet 5. The author says price and performance information was missing from the supplied context.
A small exploratory test. A candidate model appearing in the list does not demonstrate a completed integration with that model, or establish an optimal routing policy.
LimboAI (@limbopeng)
Published:
Source checked:
[22] Refactoring intent recognition with Jev and DeepSeek
LimboAI reports replacing intent-recognition work in a project with Jev alongside DeepSeek V4 Flash and describes a positive experience with speed and cost.
Qualitative personal feedback. No dataset, baseline, elapsed time or billing totals are given.
Yanhua (@yanhua1010)
Published:
Source checked:
[23] Searching train services with Pi, DeepSeek and Jev Ultrafast
Yanhua describes running Browser Use’s Jev Ultrafast with Pi and DeepSeek to look up train services on 12306 and organize the results.
Personal browser-automation demonstration. The post gives no DeepSeek version or numerical cost, and does not claim a completed ticket purchase.
Browser Use
Source checked:
[24] Jev Ultrafast action selection and text generation
The agent builds an indexed element table, asks Jev for an operation and compatible target, and calls a text model for TYPE_TEXT. The README describes an OpenAI-compatible helper configurable for Gemini, GLM and DeepSeek.
Project documentation, not a reproduction of Yanhua’s setup. The checked README defaults to Mercury 2.5, requires independent outcome checks, and lists unsupported browser features.