AI Models & Platforms

H Company Releases Holo4, Open-Weight Models for Computer-Use Agents

mm
Add Unite.AI to your preferred sources on Google

H company released Holo4, its new series of agentic computer-use models, on September 28, 2026, in 27B dense and 35B-A3B Mixture of Experts sizes, with open weights on Hugging Face and API availability. The company reports the 27B model scores 85.2% on the OSWorld benchmark at $0.08 per task.

Model Lineup and Availability

According to the company, Holo4 interacts with software through any available interface: graphical user interfaces, code, MCP and APIs. It clicks and types on a screen, writes and runs its own code, and calls MCP or API tools, selecting whichever approach fits the task. The same model runs on desktops, on the web, on Android, in a code sandbox and against business APIs, and is called the same way in each case, with no need to select a different model per platform.

Both sizes are available on the H Models API, and weights are published on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF formats. Alongside Holo4, the company also released Holotron4 Nano, an updated version of Holotron 3.

The Holo4-27B model card identifies the model as a 27B-parameter vision-language model built on the Qwen3.8 dense architecture, with a maximum context length of 262,144 tokens and target environments covering web, desktop and mobile. The card states that the weights are available under the non-commercial CC BY-NC 4.0 license and describes use with the company’s hai-agents harness, which sends screenshots and tool results to the model and executes its requested clicks, typing, code and tool calls.

The company’s newsroom post lists Holo4-35B-A3B under Apache 2.0 at $0.30 per million input tokens, $0.03 per million cached tokens and $2.00 per million output tokens, with 262K context. It lists Holo4-27B as research-only at $0.40 input, $0.04 cached and $3.00 output per million tokens, also with 262K context.

Reported Benchmarks and Cost Per Task

The company’s benchmark table reports Holo4 27B at 85.2% on OSWorld with a cost of $0.08 per task and Holo4 35B-A3B at 80.8% at $0.05, against 84.3% at $0.22 for Qwen3.8 27B, the 27B model’s base. Listed frontier scores on OSWorld are 86.0% for Fable 5, 78.7% for GPT-5.5 and 86.1% for Qwen3.8 Max.

On OSWorld 2.0, which covers long computer workflows, the company reports Holo4 27B at a 61.7% score and 41.5% success rate at $1.22 per task, and Holo4 35B-A3B at 30.9% at $0.61. The table lists Opus 5.5 at 81.8% and $8.48 per task, GPT-6 Astra at 73.5% and $9.07, and Qwen3.8 27B at 48.0% and $3.49. The company says Holo4 trails only the strongest closed models on long workflows, and that it does so with orders of magnitude fewer parameters and at much lower cost.

On AutomationBench, a business-automation benchmark, the reported scores are 45.4% at $0.05 per task for Holo4 27B and 34.5% at $0.02 for 35B-A3B. The company notes that 480 of the 600 tasks in the public v1.0.6 set fall within the split from which it collected training data; on the 120 held-out tasks, it reports 49.3% for the 27B model and 31.7% for the MoE model. It says it will report private-set results once Holo4 is evaluated on the official private set.

The table also reports scores of 44.1% for the 27B model and 30.9% for the MoE on ALE-CLI, the 105-task Linux split of the Agents’ Last Exam benchmark, and AndroidWorld scores of 85.1% and 77.6%. On internal Agentic Task Factory evaluation tasks that the company says were not used in training, the 27B model scores 80.2% on web tasks, 89.4% on MCP and 72.0% on desktop.

The company states that Holo4’s figures come from its own harness, as a mean over two to four runs with a single run on OSWorld 2.0 and ALE-CLI, while comparison scores come from model cards, official leaderboards and providers’ charts across differing harnesses and effort levels. It has open-sourced every trajectory behind its public benchmark scores, replayable through a dedicated viewer and downloadable as a Hugging Face dataset.

Training Pipeline, Holotron4 Nano and Next Step

Holo4 was trained through supervised fine-tuning and reinforcement learning on environments and tasks including those generated by the company’s Agentic Task Factory, an internal set of pipelines that builds interactive environments and verifiable tasks from documentation alone, such as screenshots of real websites or open-source software. The company says the factory has produced about 10,000 tasks so far, roughly 4,000 of them for web apps and about 3,000 each for MCP servers and desktop environments, including hybrid environments that expose the same state through a GUI and MCP.

Supervised fine-tuning used 127B tokens, about three quarters of them successful agentic trajectories split across desktop (45%), web (14%), MCP and API (12%) and mobile (3%), with the remainder covering multimodal reasoning, GUI grounding and text-only tool use and coding. Asynchronous online reinforcement learning on long-horizon tasks then trained two specialized LoRA experts, one for desktop and web and one for terminal, MCP and API, which were merged back into the fine-tuned model with equal weight and no further training.

The company also rebuilt its harness, the loop that executes the model’s actions and manages its context over hundreds of steps, using feedback from agentic performance on OSWorld 2.0. In that process, agents tagged why each task failed and engineers reviewed their fixes; the largest changes were a reliable memory that can keep track of hundreds of steps and a shell on the desktop machine itself.

As a member of the NVIDIA Nemotron Coalition, H company applied the same post-training stack to Nemotron 3 Nano Omni to produce Holotron4 Nano. The company reports absolute percentage-point gains over the base model on five benchmarks: OSWorld from 21.0 to 76.3, OSWorld 2.0 from 0.2 to 7.9, AutomationBench from 19.4 to 35.6, PinchBench from 84.7 to 88.6 and ALE Linux from 0.6 to 8.5. It says the gains show that its recipe transfers across foundation models and is not size-specific.

The company said it will release optimized DSpark drafter checkpoints in the days following the announcement to further accelerate inference.

Jonas Reeve is an AI-generated analyst at Unite.AI, focusing on cognitive AI, artificial general intelligence (AGI), and the theoretical foundations of machine intelligence. His work explores how learning, reasoning, memory, and abstraction emerge in both biological and artificial systems, drawing connections between modern AI architectures and long-standing questions in cognitive science and philosophy of mind.

With a conceptual and reflective approach, Jonas examines frameworks such as reasoning models, agentic systems, emergent cognition, and alignment theory, aiming to clarify what progress toward AGI actually means—and what it does not. Rather than chasing timelines or hype, he emphasizes first principles, conceptual rigor, and the limits of current models.

Articles authored by Jonas Reeve are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, clarity, and responsible discussion of advanced AI concepts.