SYSTEM ONE MODEL INTELLIGENCE DIRECTORY|Community workflows and reports
Independent site
Last updated:
TRANSPARENT PROVENANCE

Sources behind the Jev reports

We read the original X posts, follow-up discussions and linked project documentation before publishing. These cards are editorial summaries, not quotations. A source check confirms what the author reported; it does not mean we reproduced the experiment. Measurements retain their task, model and sample limits.

Source notes (24)

TypeSafe AI

Source checked:

Official

[1] Jev and the Choice, Score and Noul primitives

The documentation describes Jev as a System One model that evaluates state and typed questions. Choice selects an option, Score evaluates a rubric, and Noul returns a probability for a statement.

API behavior documented by the provider. This page does not establish a universal latency, accuracy or cost advantage.

TypeSafe AI

Source checked:

Official

[2] Confidence, probabilities and fallback decisions

Choice and Score include confidence derived from their probability distributions. Noul does not have a separate confidence field. TypeSafe recommends choosing thresholds for the domain and consequences of an action.

Implementation reference. A confidence value is not a guarantee that an individual decision is correct.

paulwei (@coolish)

Published:

Source checked:

X.com

[3] Trying Jev after GPT-6 Astra in Slay the Spire 2

The author describes GPT-6 Astra as capable but slow in an earlier game run, then reports roughly 0.7 seconds of action deliberation with Jev.

Personal game experiment with a video. The timing is not a controlled, repeated GPT-versus-Jev benchmark.

Jev / GPTRead original

paulwei (@coolish)

Published:

Source checked:

X.com

[4] Astra revises Jev prompts after a failed game run

The first Ascension 10 attempt reached floor 6. The author says Astra then iterated on Jev prompts and the pair reached floor 17, while describing Jev as weaker at the game than Astra.

A follow-up from the same experiment, not independent corroboration. Reaching floor 17 is not a reported completed run.

Jev / GPTRead original

Sac (@Saccc_c)

Published:

Source checked:

X.com

[5] Codex and Jev for adding a Mac calendar event

Sac describes a Codex + Jev computer-use experiment and compares adding the same Mac calendar event. The reported experience is smoother, with similar token consumption.

Qualitative report and comparison video. The post does not identify the underlying GPT version or give a repeatable latency measurement.

Jev / GPT / CodexRead original

tamara (@tamarajtran)

Published:

Source checked:

X.com

[6] Scoring tool calls for context compaction

The project author proposes using Jev to score tool calls and discard irrelevant material during compaction.

Project announcement. The repository says its native demo app is scripted and makes no API calls, so its animation is not timing evidence.

Jev / ClaudeRead original

Alex Volkov (@altryne)

Published:

Source checked:

X.com

[7] A user report of Claude context reduction

Volkov reports reducing a Claude session from nearly one million tokens to about 86,000 in roughly one second using fast-jev-compaction.

Individual experience, not a controlled benchmark or a measure of downstream task accuracy. The timestamp is the edited-post time shown by X.

Jev / ClaudeRead original

tamaratran

Source checked:

GitHub

[8] fast-jev-compaction implementation and limitations

The plugin scores tool calls and results, preserves user and assistant text in its output, and keeps calls paired with their results. Its hook falls back to built-in compaction on errors or insufficient reduction.

The README documents early-access Claude Code function hooks, estimated token sizes and a scripted native demo. It explicitly warns that a probability does not prove a result is safe to delete.

Jev / ClaudeRead original

nikhil mudholkar (@nikhilmudholkar)

Published:

Source checked:

X.com

[9] Business-email classification against Gemini

Mudholkar compares Jev with Gemini on 1,565 German and English business emails in 10 categories. Gemini performs better on overall accuracy; the author is interested in Jev for its uncertainty signal.

An author-reported benchmark. It compares separate models, rather than measuring a combined Jev + Gemini pipeline.

Jev / GeminiRead original

nikhil mudholkar (@nikhilmudholkar)

Published:

Source checked:

X.com

[10] Email dataset composition and shared instructions

The dataset contains 1,201 real emails and 364 AI-written examples for rare categories, mostly in German. The three models receive the same email text and category instructions.

Methodology from the benchmark thread. The dataset is not entirely real correspondence.

Jev / GeminiRead original

nikhil mudholkar (@nikhilmudholkar)

Published:

Source checked:

X.com

[11] Email accuracy results for Jev and two Gemini models

Reported overall accuracy is 96.4% for Jev, 97.5% for Gemini 3.5 Flash-Lite and 98.5% for Gemini 3.8 Flash. Equal weighting of categories yields 92.0%, 94.6% and 96.9%, respectively.

Results against the author’s reference labels in this dataset. These are not general accuracy scores for the models.

Jev / GeminiRead original

nikhil mudholkar (@nikhilmudholkar)

Published:

Source checked:

X.com

[12] Where the email-classification mistakes occurred

The author reports that 737 predictions at 99% confidence or above matched the reference labels. Below 70% confidence, nearly half were wrong.

An observed relationship in one dataset, not a universal threshold or proof of error-free decisions.

Jev / GeminiRead original

nikhil mudholkar (@nikhilmudholkar)

Published:

Source checked:

X.com

[13] Reported cost per 1,000 business emails

The author reports $0.08 for Jev, $0.80 for Gemini 3.5 Flash-Lite and $1.79 for Gemini 3.8 Flash per 1,000 emails in these runs.

Workload-specific costs, not provider pricing or the cost of the whole 1,565-email dataset. The author notes that savings are small at ordinary inbox volumes.

Jev / GeminiRead original

nikhil mudholkar (@nikhilmudholkar)

Published:

Source checked:

X.com

[14] Attachments and synthetic examples in the email test

Attachments were excluded, and 364 emails were generated to match chosen labels. The author describes the lack of non-text input as a barrier to production use.

Limitations supplied by the experiment’s author. This source does not demonstrate processing email attachments.

Jev / GeminiRead original

Rahul Kumar (@rahulbuildsmore)

Published:

Source checked:

X.com

[15] Jev and Gemini labeling 1,000 app reviews

Kumar compares sentiment, topic, bug and churn labels on 1,000 app reviews. He reports Jev at 4.6 seconds and $0.023, versus Gemini 3.8 Flash at 18.8 seconds and $0.158.

The post links to a public project with recorded runs and measurement notes. It is a side-by-side comparison, not a combined workflow.

Jev / GeminiRead original

goodrahstar

Source checked:

GitHub

[16] Column Race recorded runs and evaluation methodology

Both lanes process 20 reviews per request at concurrency 8. The recorded pair uses Jev 1.13.0 and Gemini 3.8 Flash. Sentiment correlation with star ratings is 0.803 and 0.822; topic agreement is 84.4% and bug agreement is 92.6%.

One recorded pair over 1,000 Android reviews. Stars are a weak sentiment proxy; topic and bug labels have no human ground truth. Agreement is not accuracy. Costs use token counts and list prices, not invoices.

Jev / GeminiRead original

Dan McAteer (@daniel_mac8)

Published:

Source checked:

X.com

[17] Ranking X posts with Jev and Grok Bot

McAteer describes retrieving 1,000 agent-related posts with the X API, using Jev for pairwise selection, and having Grok Bot surface five tips.

A first-hand workflow description. No model version, comparison count, latency or cost breakdown is supplied.

Jev / GrokRead original

codila (@0xCodila)

Published:

Source checked:

X.com

[18] A Jev router in front of GrokBot actions

The author describes consulting Jev before browser, research and retry actions, starting in shadow mode, inspecting logs, and retaining a bypass and human control over irreversible actions.

A personal setup report with promotional language. Its broad performance claims and seven-minute setup claim are not treated as measurements.

Jev / GrokRead original

Hassan (@nutlope)

Published:

Source checked:

X.com

[19] Jev first-pass email screening with Kimi K3 fallback

Hassan tests 50 legitimate and 50 fraudulent emails. Jev classifies them in 1.42 seconds; 31 cases below a 95% confidence threshold go to Kimi K3. The full pipeline takes 16 seconds, gets 96 of 100 correct and costs about $0.07.

Author-reported demonstration, not independently rerun here. The stated cost split is $0.003 for Jev and $0.068 for Kimi. The 1.42-second figure excludes the Kimi stage.

Jev / KimiRead original

yonsakhan (@yonsakhan)

Published:

Source checked:

X.com

[20] Using Kimi K3 to develop a Jev-powered comment filter

The author says they developed an X comment-filtering userscript with Kimi K3 and use Jev for semantic spam decisions, including obfuscated text. They describe reversible collapsing and a local cache.

A development report. Kimi helped write the software; the post does not establish that Kimi participates in its runtime classification loop.

Jev / KimiRead original

Jing Wang (@jingwangtalk)

Published:

Source checked:

X.com

[21] A model-routing experiment sensitive to the candidate list

Wang experiments with a Jev router over DeepSeek, GLM, Kimi, GPT and Claude candidates. Adding Claude Opus 5 changes the selected model from DeepSeek Flash to Claude Sonnet 5. The author says price and performance information was missing from the supplied context.

A small exploratory test. A candidate model appearing in the list does not demonstrate a completed integration with that model, or establish an optimal routing policy.

Jev / DeepSeek / Kimi / GPT / ClaudeRead original

LimboAI (@limbopeng)

Published:

Source checked:

X.com

[22] Refactoring intent recognition with Jev and DeepSeek

LimboAI reports replacing intent-recognition work in a project with Jev alongside DeepSeek V4 Flash and describes a positive experience with speed and cost.

Qualitative personal feedback. No dataset, baseline, elapsed time or billing totals are given.

Jev / DeepSeekRead original

Yanhua (@yanhua1010)

Published:

Source checked:

X.com

[23] Searching train services with Pi, DeepSeek and Jev Ultrafast

Yanhua describes running Browser Use’s Jev Ultrafast with Pi and DeepSeek to look up train services on 12306 and organize the results.

Personal browser-automation demonstration. The post gives no DeepSeek version or numerical cost, and does not claim a completed ticket purchase.

Jev / DeepSeekRead original

Browser Use

Source checked:

GitHub

[24] Jev Ultrafast action selection and text generation

The agent builds an indexed element table, asks Jev for an operation and compatible target, and calls a text model for TYPE_TEXT. The README describes an OpenAI-compatible helper configurable for Gemini, GLM and DeepSeek.

Project documentation, not a reproduction of Yanhua’s setup. The checked README defaults to Mercury 2.5, requires independent outcome checks, and lists unsupported browser features.

Jev / DeepSeek / GeminiRead original