AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

H Company says its new Holo4 models can switch among screen interaction, code execution and MCP or API tools within a single agent. The company reports a 61.7% score for Holo4 27B on OSWorld 2.0, while noting that benchmark results and cost comparisons use different evaluation setups.

H Company has announced Holo4, a series of agentic models designed to carry out software tasks through graphical interfaces, code, MCP tools and APIs. The company says the same model can use different interfaces across desktops, web apps, Android, code sandboxes and business APIs, an approach aimed at workflows that span more than one kind of software access.

The release includes a 27-billion-parameter dense model and a 35B-A3B Mixture of Experts model, both available through H Models API, according to H Company. The company also announced Holotron4 Nano, an updated version of Holotron 3. It says Holo4 was trained with supervised and reinforcement learning across a large set of environments and tasks, including tasks generated by its Agentic Task Factory.

H Company reports that Holo4 27B scored 61.7% on OSWorld 2.0, a benchmark for computer use. The report compares that result with 81.8% for Opus 5.5 and gives Holo4 35B-A3B a score of 30.9%. The company says Holo4 improves on its Qwen base models and argues that its models can perform well at lower cost, but the published comparisons combine results from different harnesses, task sets and pricing assumptions.

The company says it is releasing the model weights in FP16, FP8 and GGUF formats and publishing trajectories behind its public benchmark scores. Those trajectories can be replayed through H Company’s viewer or downloaded from Hugging Face, offering a way to inspect the steps behind reported outcomes.

At a glance
announcementWhen: Announced in the source report; publica…
The developmentH Company announced Holo4, a model series designed to perform computer-use tasks through GUIs, code, MCP and APIs.

One Model Across Business Software

Many computer-use agents are built around a particular way of interacting with software: clicking and typing through a screen, or calling tools through an API. H Company’s case for Holo4 is that a single business task can require both. A model might need to inspect a screen, write code in a sandbox and then use a business API, for example.

If the company’s account holds up in independent use, a model that can move among those interfaces could simplify how organizations build agents for mixed software environments. H Company says customers would not need to choose a different model for each platform. The release materials, however, establish the company’s design and benchmark claims, not how reliably Holo4 will handle varied live business workflows.

From Benchmarks to Workflows

H Company presents Holo4 as a successor to an earlier model and frames it around professional software tasks, rather than academic benchmarks alone. Its examples compare Holo4 27B with the Qwen 27B base model using the same prompt and harness. They include building 3D models in FreeCAD and creating a game in Godot; the report describes these as demonstrations, not independent evaluations.

The benchmark comparisons have important limits. For OSWorld 2.0, H Company says Holo4 costs are estimated from tokens used in a single run at its API rates. Cost figures for other models draw on different sources and assumptions. For AutomationBench, Holo4 and two Qwen models were tested on version 1.0.6 in H Company’s internal harness, while other models’ scores and cost figures come from public materials that use a private task set. The company says it will report Holo4 results on that private set after evaluation.

“Real work is not siloed that way, and a single business task can require combining these different approaches.”

— H Company

Independent Results Still Pending

The supplied report is H Company’s own announcement. It does not provide independent confirmation of the benchmark scores, cost advantages or performance on business workflows. The OSWorld figures draw on sources and evaluation setups that differ, and the AutomationBench comparison mixes results from a public set and a private set. Holo4’s result on AutomationBench’s private set has not yet been reported by the company.

The announcement also does not establish how often Holo4 completes lengthy real-world tasks successfully, how performance varies across the listed interfaces, or what costs look like across repeated use. The report describes example tasks but does not provide enough detail to treat them as a broad measure of reliability.

Private Benchmark Evaluation Ahead

H Company says it plans to publish Holo4’s AutomationBench result on the private task set after evaluation. It has also made benchmark trajectories available for replay and download, which gives researchers and users material to examine alongside the reported scores. The source does not specify a date for the private-set results or other evaluation milestones.

For now, the models are available through H Models API, with weights offered in several formats, according to the announcement. Further evidence from independent testing and real deployments will help show whether Holo4’s cross-interface design translates into consistent performance and the cost savings H Company claims.

Key Questions

What is Holo4?

Holo4 is H Company’s new series of agentic models for software tasks. The company says the models can interact with screens, write and run code, and call MCP or API tools.

Which Holo4 models did H Company announce?

The announcement lists a 27B dense model and a 35B-A3B Mixture of Experts model. It also introduces Holotron4 Nano, an updated version of Holotron 3.

How did Holo4 perform on OSWorld 2.0?

H Company reports a 61.7% score for Holo4 27B and 30.9% for Holo4 35B-A3B. It compares the 27B score with 81.8% for Opus 5.5; the announcement notes that evaluation and cost comparisons involve differing sources and setups.

Are Holo4’s cost claims independently verified?

The source is H Company’s report, and it does not provide independent verification of the claimed cost advantage. Its comparisons use different harnesses, task sets and pricing assumptions.

What Holo4 results are still to come?

H Company says it will report Holo4’s result on AutomationBench’s private set after evaluation. It has not given a date for that result in the supplied report.

Source: rss

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Ultimate Guide To Using AI For Dynamic Data In Sheets Canvas

Google introduces Sheets Canvas, a Gemini-powered feature enabling interactive dashboards from spreadsheet data via natural-language prompts, now rolling out globally.

Tracking Food Trends: Lahori Zeera Gains Global Attention

GDELT’s media monitoring gave the Lahori Zeera trend an 88/100 signal score as global coverage of the Pakistani drink surges. What is confirmed, and what isn’t.

How Stony Brook Researchers Designed A Blueprint For Self-Improving AI

A headline reports that Stony Brook researchers developed a blueprint for self-improving AI, but provides no study details or evidence.