Holo4’s Approach To General-Purpose Computer-Use Agents
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Holo4’s Approach To General-Purpose Computer-Use Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

H Company has released Holo4, a series of open-weight agentic models in 27B dense and 35B-A3B MoE sizes, designed to operate software through GUIs, code, MCP and APIs with a single model. The 27B version scores a self-reported 61.7% on OSWorld 2.0, trailing the strongest closed model, with all benchmark trajectories published for public audit.

H Company has released Holo4, a new series of open-weight agentic models built to operate software through graphical interfaces, code, MCP and APIs using a single model, as detailed in the original analysis. The series ships in two sizes — a 27B dense model and a 35B-A3B Mixture of Experts model — available on the H Models API and for download on Hugging Face. According to the company, the 27B version scores 61.7% on OSWorld 2.0, trailing only the strongest closed models while using far fewer parameters and lower cost.

Holo4 is positioned as a general-purpose computer-use agent: it clicks and types on screens, writes and runs its own code, and calls MCP or API tools, selecting whichever interface suits the task. H Company says the same model is invoked the same way across desktops, the web, Android, a code sandbox and business APIs. The company argues that most existing agentic models are trained for a single interface — GUI-focused models fail without a screen, while tool-calling models cannot handle applications that lack APIs.

The models were built on a Qwen base — Qwen3.8 27B for the dense model and Qwen3.6 35B-A3B for the MoE variant, according to the company’s benchmark notes — and trained through supervised and reinforcement learning across a large set of environments and tasks, including tasks generated by H Company’s Agentic Task Factory. The company reports substantial improvement over the Qwen base, demonstrated in side-by-side examples on professional software such as FreeCAD 3D modeling and Godot game design, run with the same prompt and harness.

On benchmarks, H Company reports Holo4 27B scoring 61.7% and Holo4 35B-A3B scoring 30.9% on OSWorld 2.0, compared with 81.8% for Opus 5.5, the strongest closed model in the comparison. The company has open-sourced every trajectory behind its public benchmark scores, viewable at trajectories.hcompany.ai and downloadable from Hugging Face. The release arrives alongside an updated Holotron 3, called Holotron4 Nano.

At a glance
announcementWhen: announced and released; evaluations ong…
The developmentH Company announced and released Holo4, two open-weight computer-use agent models, along with full benchmark trajectory data for public verification.
At a glance
announcementWhen: announced via Hugging Face and company…
The developmentH Company announced the release of Holo4, a two-model series of open-weight computer-use agents, along with an updated Holotron4 Nano and open-sourced benchmark trajectories.

Open-Weight Agents Closing the Gap

The release matters because open-weight computer-use agents remain rare at this reported performance level. If Holo4’s scores hold up under independent evaluation, businesses could run capable software-automation agents at a fraction of the cost of frontier closed models, with the flexibility of self-hosting.

The claimed cost-performance gap — 61.7% on OSWorld 2.0 from a 27B model against 81.8% from a much larger closed model — would represent meaningful progress for smaller, cheaper agents. The multi-interface design also addresses a practical limitation: real business tasks often mix screen work, code and API calls, and single-interface models break at those boundaries.

H Company’s decision to publish all benchmark trajectories lets outside parties verify each step, a level of transparency most closed-model providers do not offer.

From Holo1 to Holo4 and Backed by Qwen

Holo4 builds on H Company’s previous agentic model line, following earlier releases in the Holo series, and arrives together with Holotron4 Nano, an updated version of its Holotron 3 system.

The company’s cost comparisons rest on specific assumptions it has disclosed: Holo4 is priced at H Models API rates for a single run, Qwen costs are calculated at Alibaba Cloud list prices (with cache hits at 20% of input price for the MoE model), and GPT and Opus effort sweeps come from OpenAI launch data. H Company cautions that releases, harnesses and task subsets differ across the compared models.

“Real work is not siloed that way, and a single business task can require combining these different approaches.”

— H Company, announcement

Claims Awaiting Independent Verification

All headline benchmark numbers are self-reported by H Company and measured in the company’s own harness, which it acknowledges differs from other models’ releases, harnesses and task subsets.

On AutomationBench, other models’ scores come from the public set while costs come from a leaderboard running the private set — a mismatch the company itself acknowledges. Holo4 has not yet been evaluated on the AutomationBench private set.

The steep score difference between the 27B dense model (61.7%) and the larger 35B-A3B MoE model (30.9%) on OSWorld 2.0 is not explained in the announcement. Real-world reliability on business workflows, beyond curated demo examples, also remains unverified by third parties.

Evaluations and Adoption Watchpoints

H Company says it will report Holo4 results on the AutomationBench private set once that evaluation is complete. Independent benchmark submissions and third-party reproductions — now possible because trajectories and weights are public — will be the next test of the company’s claims.

Developers can access the models through the H Models API or download the full collection from Hugging Face in FP16, FP8 and GGUF formats.

Key Questions

What is Holo4?

Holo4 is H Company’s new series of open-weight agentic models designed to operate software through GUIs, code, MCP and APIs using a single model. It comes in a 27B dense version and a 35B-A3B Mixture of Experts version.

How does Holo4 perform on benchmarks?

According to H Company, Holo4 27B scores 61.7% on OSWorld 2.0 and the 35B-A3B version scores 30.9%, compared with 81.8% for Opus 5.5, the strongest closed model. These figures are self-reported and measured in the company’s own harness.

Are the benchmark results independently verified?

No. All headline numbers are self-reported by H Company. However, the company has published every trajectory behind its public benchmark scores at trajectories.hcompany.ai, allowing outside parties to audit and attempt reproduction.

How can developers access Holo4?

The models are available through the H Models API and for download on Hugging Face in FP16, FP8 and GGUF formats.

Why does the smaller model score higher than the larger one?

The announcement does not explain why the 27B dense model (61.7%) outscores the larger 35B-A3B MoE model (30.9%) on OSWorld 2.0. This remains an open question.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What ByteDance’s Founder Says About AI Distillation And Its Risks

ByteDance’s founder reportedly instructed employees to avoid AI distillation, signaling potential shifts in AI development strategies amid uncertain details.

SenseTime’s Breakthrough: Turning AI Investments Into Its First Profitable Half-Year

SenseTime announced a RMB 620 million profit for the first half, marking its first profit since listing, though key details remain undisclosed.

Game 1 Total Kills Prediction: Is It Going Over Or Under 65.5?

Analysis of the upcoming Dota 2 match between 1win and Team Nemesis, focusing on whether total kills in Game 1 will go over or under 65.5, with key details and uncertainties.

Game 2: Any Player Penta Kill?

Search interest spikes as speculation grows over a potential pentakill in Game 2, with a new Polymarket market indicating a 50% chance. Details remain unconfirmed.