AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What Can Jev Do For AI Decisions? 24 Ways To Put It To Work on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Thorsten Meyer’s Sept. 29 article describes 24 potential uses for Jev, a tool that returns typed answers for software decisions. He says three uses are live in his publishing operation, with about 90,000 decisions made so far; 12 more meet his proposed fit test, while seven need measurement and two are poor fits. The results and performance figures come from Meyer, and the supplied source ends before detailing all 24 use cases.

Thorsten Meyer published a guide on Sept. 29 describing 24 potential uses for Jev, a tool that returns structured answers to questions about text or JSON so software can make decisions. Meyer says three applications are already live in his publishing operation, which has processed about 90,000 decisions, while the remaining ideas range from strong candidates to cases that need measurement or are poor fits.

Meyer says Jev accepts a state and typed questions in a single call, returning answers without prose for code to interpret. The answer types include a yes-or-no probability, a choice among options with probabilities and confidence, or a score on ordered levels. He reports a call takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens. Those performance and cost figures are claims in his article; the supplied material does not include independent testing details.

The three live publishing uses are a relevance gate for matching stories to a site, an English-language check, and a fallback topic classifier. Meyer reports scanning 78,889 articles for $2.01 with the language check, finding 1,576 non-English items and fixing 1,553. For the classifier, he reports 89% agreement with a frontier large language model overall, and 97% to 99% agreement for answers with confidence of at least 0.8. The source does not specify the evaluation sample behind that comparison.

His inventory assigns 12 uses a “strong fit”, seven “measure first,” and two “poor fit,” alongside the three live applications. Among the publishing examples described are disclosure checks and comment moderation, which he rates as strong fits; thin-source detection, product matching in roundups, and headline quality are marked for measurement. Same-event deduplication is a poor fit in his account because a canary test found no duplicates.

At a glance
reportWhen: Published Sept. 29, 2026
The developmentThorsten Meyer published a guide outlining 24 ways to use Jev for automated decisions, reporting three live publishing applications and setting out a four-part test for deciding where the tool fits.

24 use cases for Jev at a glance

Publishing, commerce, software, business operations and the home, sorted by fit.

Every use case, coloured by how well it fits

Start in the green. Amber needs a measurement first. Red fails at least one of the four conditions.
livestrong fitmeasure firstpoor fit

Proven in production

1Relevance gate: story and site2Language check3Classifier fallback

Publishing and content

4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderation

Commerce and support

10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triage

Software and AI systems

15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triage

Business ops and home

21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent

15 of 24 are ready to build or already running

3
12
7
2
Live
Strong fit
Measure first
Poor fit
Live: in my fleet today. Strong fit: meets high volume, narrow question, cheap errors and a visibly failing heuristic. Measure first: the failing heuristic is unproven.
From “24 Ways to Use Jev” on thorstenmeyerai.com. Figures are my own production measurements, September 2026, rounded, unless marked illustrative.

Where Automated Checks May Help

The guide’s practical argument is that small, repeated decisions can be automated when mistakes are inexpensive or uncertain cases can be sent to a person or a more capable system. That could let publishers apply checks across far more items than they would review manually, while keeping higher-stakes or ambiguous calls out of automatic handling.

Meyer’s proposed confidence threshold is central to that approach: software acts on clear answers and routes the gray area elsewhere. He says Jev agreed with a frontier model 97% to 99% of the time at confidence of 0.8 or higher in a 31-topic classification test, compared with 42% below 0.5. These are author-reported measurements, not an independently verified benchmark in the provided material. The suggested uses therefore matter as a deployment framework, but the reported accuracy should not be generalized without further evaluation.

A Four-Part Test for Jev

Meyer advises using Jev only when a task has high volume, asks a narrow question rather than requiring multi-step reasoning, has cheap errors or a safe route for uncertain answers, and has a heuristic that has been shown to fail. He cautions against adding an AI check just because it is inexpensive: if a keyword rule already works, it should stay in place until evidence shows a problem.

For proposed uses, Meyer recommends replaying 300 to 500 past decisions, comparing results overall and by confidence band, and reviewing 20 disagreements to determine which system was right. He says to integrate the tool only when the high-confidence band reaches 95%, then use a separate feature flag, begin with a 5% to 10% canary, and expand from there. The article presents these as his operating recommendations.

The supplied source describes only part of the 24-use inventory. It covers three live applications and several publishing checks, then cuts off as the commerce and customer-operations section begins. The full set of proposed applications across software, business operations, and the home cannot be confirmed from the material provided.

““Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.””

— Thorsten Meyer

Evidence Behind the Reported Results

The figures in the article are reported by Meyer. The supplied source does not describe an independent audit, provide the underlying data, or explain in detail how the comparison with a frontier model was conducted. It also does not identify the model, evaluation date, or sample size for that comparison beyond noting a 31-topic classification.

The source excerpt does not include the full 24-item list, so most proposed uses and their fit ratings cannot be assessed here. Meyer says seven ideas need measurement because a visibly failing heuristic has not been established, but the excerpt does not name all seven. It is also unclear how results from his publishing operation would transfer to other organizations, content types, or decision risks.

Testing Before Wider Rollout

Meyer’s next step for each proposed application is to test it against real past decisions, inspect disagreements, and confirm that high-confidence answers meet his stated accuracy bar. He recommends enabling a successful check behind a separate flag, starting with a 5% to 10% canary before wider use.

The article excerpt gives no schedule for further deployments or independent validation. Readers considering a use case would need the complete inventory and their own measured results to judge whether the tool fits their workload.

Key Questions

What does Jev do?

According to Meyer, Jev takes text or JSON plus typed questions and returns structured answers, such as a probability, a choice among categories, or a score. Software can use those answers to route or filter work.

How many Jev applications does Meyer say are live?

Meyer reports three live uses in his publishing operation: story relevance, English-language checking, and fallback topic classification. He says they have handled about 90,000 decisions in total.

Are the accuracy and cost figures independently verified?

The supplied article presents these as Meyer’s measurements and estimates. It does not provide independent verification or enough methodology to reproduce the reported model-agreement figures.

What kinds of tasks does Meyer say are a good fit?

He recommends high-volume tasks with narrow questions, low-cost errors or a safe route for uncertain cases, and an existing heuristic that has been shown to fail. He advises measuring performance before connecting the tool to live decisions.

Does the source list all 24 uses?

No. The supplied excerpt covers the live publishing uses and part of the publishing section, then ends as the commerce and customer-operations section starts. The remaining use cases and ratings are not available in the provided material.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Codex In ChatGPT Desktop App For Linux Is Now In Preview

OpenAI’s ChatGPT Linux desktop app now includes Codex in preview, enhancing coding capabilities for users. The update is currently in testing.

AI Tools and Automation: The Essential Guide to Working Smarter

AIThis post was created with the assistance of artificial intelligence (AI).Artificial intelligence…

DeepSeek V4 Pro 0813

DeepSeek has announced the release of V4 Pro 0813, a new version of its AI-powered search platform, with confirmed improvements and upcoming features.