🔍 Read the full analysis: What Can Jev Do For AI Decisions? 24 Ways To Put It To Work on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Thorsten Meyer’s Sept. 29 article describes 24 potential uses for Jev, a tool that returns typed answers for software decisions. He says three uses are live in his publishing operation, with about 90,000 decisions made so far; 12 more meet his proposed fit test, while seven need measurement and two are poor fits. The results and performance figures come from Meyer, and the supplied source ends before detailing all 24 use cases.
Thorsten Meyer published a guide on Sept. 29 describing 24 potential uses for Jev, a tool that returns structured answers to questions about text or JSON so software can make decisions. Meyer says three applications are already live in his publishing operation, which has processed about 90,000 decisions, while the remaining ideas range from strong candidates to cases that need measurement or are poor fits.
Meyer says Jev accepts a state and typed questions in a single call, returning answers without prose for code to interpret. The answer types include a yes-or-no probability, a choice among options with probabilities and confidence, or a score on ordered levels. He reports a call takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens. Those performance and cost figures are claims in his article; the supplied material does not include independent testing details.
The three live publishing uses are a relevance gate for matching stories to a site, an English-language check, and a fallback topic classifier. Meyer reports scanning 78,889 articles for $2.01 with the language check, finding 1,576 non-English items and fixing 1,553. For the classifier, he reports 89% agreement with a frontier large language model overall, and 97% to 99% agreement for answers with confidence of at least 0.8. The source does not specify the evaluation sample behind that comparison.
His inventory assigns 12 uses a “strong fit”, seven “measure first,” and two “poor fit,” alongside the three live applications. Among the publishing examples described are disclosure checks and comment moderation, which he rates as strong fits; thin-source detection, product matching in roundups, and headline quality are marked for measurement. Same-event deduplication is a poor fit in his account because a canary test found no duplicates.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Automated Checks May Help
The guide’s practical argument is that small, repeated decisions can be automated when mistakes are inexpensive or uncertain cases can be sent to a person or a more capable system. That could let publishers apply checks across far more items than they would review manually, while keeping higher-stakes or ambiguous calls out of automatic handling.
Meyer’s proposed confidence threshold is central to that approach: software acts on clear answers and routes the gray area elsewhere. He says Jev agreed with a frontier model 97% to 99% of the time at confidence of 0.8 or higher in a 31-topic classification test, compared with 42% below 0.5. These are author-reported measurements, not an independently verified benchmark in the provided material. The suggested uses therefore matter as a deployment framework, but the reported accuracy should not be generalized without further evaluation.
A Four-Part Test for Jev
Meyer advises using Jev only when a task has high volume, asks a narrow question rather than requiring multi-step reasoning, has cheap errors or a safe route for uncertain answers, and has a heuristic that has been shown to fail. He cautions against adding an AI check just because it is inexpensive: if a keyword rule already works, it should stay in place until evidence shows a problem.
For proposed uses, Meyer recommends replaying 300 to 500 past decisions, comparing results overall and by confidence band, and reviewing 20 disagreements to determine which system was right. He says to integrate the tool only when the high-confidence band reaches 95%, then use a separate feature flag, begin with a 5% to 10% canary, and expand from there. The article presents these as his operating recommendations.
The supplied source describes only part of the 24-use inventory. It covers three live applications and several publishing checks, then cuts off as the commerce and customer-operations section begins. The full set of proposed applications across software, business operations, and the home cannot be confirmed from the material provided.
““Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.””
— Thorsten Meyer
Evidence Behind the Reported Results
The figures in the article are reported by Meyer. The supplied source does not describe an independent audit, provide the underlying data, or explain in detail how the comparison with a frontier model was conducted. It also does not identify the model, evaluation date, or sample size for that comparison beyond noting a 31-topic classification.
The source excerpt does not include the full 24-item list, so most proposed uses and their fit ratings cannot be assessed here. Meyer says seven ideas need measurement because a visibly failing heuristic has not been established, but the excerpt does not name all seven. It is also unclear how results from his publishing operation would transfer to other organizations, content types, or decision risks.
Testing Before Wider Rollout
Meyer’s next step for each proposed application is to test it against real past decisions, inspect disagreements, and confirm that high-confidence answers meet his stated accuracy bar. He recommends enabling a successful check behind a separate flag, starting with a 5% to 10% canary before wider use.
The article excerpt gives no schedule for further deployments or independent validation. Readers considering a use case would need the complete inventory and their own measured results to judge whether the tool fits their workload.
Key Questions
What does Jev do?
According to Meyer, Jev takes text or JSON plus typed questions and returns structured answers, such as a probability, a choice among categories, or a score. Software can use those answers to route or filter work.
How many Jev applications does Meyer say are live?
Meyer reports three live uses in his publishing operation: story relevance, English-language checking, and fallback topic classification. He says they have handled about 90,000 decisions in total.
Are the accuracy and cost figures independently verified?
The supplied article presents these as Meyer’s measurements and estimates. It does not provide independent verification or enough methodology to reproduce the reported model-agreement figures.
What kinds of tasks does Meyer say are a good fit?
He recommends high-volume tasks with narrow questions, low-cost errors or a safe route for uncertain cases, and an existing heuristic that has been shown to fail. He advises measuring performance before connecting the tool to live decisions.
Does the source list all 24 uses?
No. The supplied excerpt covers the live publishing uses and part of the publishing section, then ends as the commerce and customer-operations section starts. The remaining use cases and ratings are not available in the provided material.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
