🔍 Read the full analysis: Holo4’s Approach To General-Purpose Computer-Use Agents on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
H Company has released Holo4, a series of open-weight agentic models in 27B dense and 35B-A3B MoE sizes, designed to operate software through GUIs, code, MCP and APIs with a single model. The 27B version scores a self-reported 61.7% on OSWorld 2.0, trailing the strongest closed model, with all benchmark trajectories published for public audit.
H Company has released Holo4, a new series of open-weight agentic models built to operate software through graphical interfaces, code, MCP and APIs using a single model, as detailed in the original analysis. The series ships in two sizes — a 27B dense model and a 35B-A3B Mixture of Experts model — available on the H Models API and for download on Hugging Face. According to the company, the 27B version scores 61.7% on OSWorld 2.0, trailing only the strongest closed models while using far fewer parameters and lower cost.
Holo4 is positioned as a general-purpose computer-use agent: it clicks and types on screens, writes and runs its own code, and calls MCP or API tools, selecting whichever interface suits the task. H Company says the same model is invoked the same way across desktops, the web, Android, a code sandbox and business APIs. The company argues that most existing agentic models are trained for a single interface — GUI-focused models fail without a screen, while tool-calling models cannot handle applications that lack APIs.
The models were built on a Qwen base — Qwen3.8 27B for the dense model and Qwen3.6 35B-A3B for the MoE variant, according to the company’s benchmark notes — and trained through supervised and reinforcement learning across a large set of environments and tasks, including tasks generated by H Company’s Agentic Task Factory. The company reports substantial improvement over the Qwen base, demonstrated in side-by-side examples on professional software such as FreeCAD 3D modeling and Godot game design, run with the same prompt and harness.
On benchmarks, H Company reports Holo4 27B scoring 61.7% and Holo4 35B-A3B scoring 30.9% on OSWorld 2.0, compared with 81.8% for Opus 5.5, the strongest closed model in the comparison. The company has open-sourced every trajectory behind its public benchmark scores, viewable at trajectories.hcompany.ai and downloadable from Hugging Face. The release arrives alongside an updated Holotron 3, called Holotron4 Nano.
Open-Weight Agents Closing the Gap
The release matters because open-weight computer-use agents remain rare at this reported performance level. If Holo4’s scores hold up under independent evaluation, businesses could run capable software-automation agents at a fraction of the cost of frontier closed models, with the flexibility of self-hosting.
The claimed cost-performance gap — 61.7% on OSWorld 2.0 from a 27B model against 81.8% from a much larger closed model — would represent meaningful progress for smaller, cheaper agents. The multi-interface design also addresses a practical limitation: real business tasks often mix screen work, code and API calls, and single-interface models break at those boundaries.
H Company’s decision to publish all benchmark trajectories lets outside parties verify each step, a level of transparency most closed-model providers do not offer.
From Holo1 to Holo4 and Backed by Qwen
Holo4 builds on H Company’s previous agentic model line, following earlier releases in the Holo series, and arrives together with Holotron4 Nano, an updated version of its Holotron 3 system.
The company’s cost comparisons rest on specific assumptions it has disclosed: Holo4 is priced at H Models API rates for a single run, Qwen costs are calculated at Alibaba Cloud list prices (with cache hits at 20% of input price for the MoE model), and GPT and Opus effort sweeps come from OpenAI launch data. H Company cautions that releases, harnesses and task subsets differ across the compared models.
“Real work is not siloed that way, and a single business task can require combining these different approaches.”
— H Company, announcement
Claims Awaiting Independent Verification
All headline benchmark numbers are self-reported by H Company and measured in the company’s own harness, which it acknowledges differs from other models’ releases, harnesses and task subsets.
On AutomationBench, other models’ scores come from the public set while costs come from a leaderboard running the private set — a mismatch the company itself acknowledges. Holo4 has not yet been evaluated on the AutomationBench private set.
The steep score difference between the 27B dense model (61.7%) and the larger 35B-A3B MoE model (30.9%) on OSWorld 2.0 is not explained in the announcement. Real-world reliability on business workflows, beyond curated demo examples, also remains unverified by third parties.
Evaluations and Adoption Watchpoints
H Company says it will report Holo4 results on the AutomationBench private set once that evaluation is complete. Independent benchmark submissions and third-party reproductions — now possible because trajectories and weights are public — will be the next test of the company’s claims.
Developers can access the models through the H Models API or download the full collection from Hugging Face in FP16, FP8 and GGUF formats.
Key Questions
What is Holo4?
Holo4 is H Company’s new series of open-weight agentic models designed to operate software through GUIs, code, MCP and APIs using a single model. It comes in a 27B dense version and a 35B-A3B Mixture of Experts version.
How does Holo4 perform on benchmarks?
According to H Company, Holo4 27B scores 61.7% on OSWorld 2.0 and the 35B-A3B version scores 30.9%, compared with 81.8% for Opus 5.5, the strongest closed model. These figures are self-reported and measured in the company’s own harness.
Are the benchmark results independently verified?
No. All headline numbers are self-reported by H Company. However, the company has published every trajectory behind its public benchmark scores at trajectories.hcompany.ai, allowing outside parties to audit and attempt reproduction.
How can developers access Holo4?
The models are available through the H Models API and for download on Hugging Face in FP16, FP8 and GGUF formats.
Why does the smaller model score higher than the larger one?
The announcement does not explain why the 27B dense model (61.7%) outscores the larger 35B-A3B MoE model (30.9%) on OSWorld 2.0. This remains an open question.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
