WaveSpeedAI

Holo4 27B vs 35B-A3B: Which Fits Your Agent?

Holo4 27B vs 35B-A3B compares long-task reliability, task economics, and weight licenses so builders can choose an agent model.

By John5 min read
Holo4 27B vs 35B-A3B: Which Fits Your Agent?

A cheap agent run becomes expensive when it fails late and forces a restart. That is the problem behind ​Holo4 27B vs 35B-A3B​. My answer: start with 27B when long-task completion is the constraint; shortlist 35B-A3B when lower attempt cost or commercial self-hosting matters more for your workload. Do not switch models yet. Look at the workflow first.

John. This is an evidence review, not hands-on. H Company’s figures are reference points, not my results.

Quick Verdict by Agent Workload

Treat the Holo4 27B vs 35B-A3B choice as a matched-task decision. Record attempts, completions, tokens, and interventions. A cheaper request is not cheaper if it creates more restarts.

What Vendor Long-Task Evidence Supports for 27B

Vendor evidence favors Holo4 27B on extended work. In its September 28, 2026 Holo4 release report, H Company reported an OSWorld 2.0 partial score of 61.7%, a 41.5% success rate, and $1.22 per task for 27B. It reported 30.9%, 12.3%, and $0.61 respectively for 35B-A3B.

Those were single runs in H Company’s harness, not independent results. The harness managed memory over hundreds of steps and provided a desktop shell. The evidence supports testing 27B first for long workflows; it does not prove the same advantage with your tools, prompts, or recovery logic.

When 35B-A3B Licensing Changes the Deployment Choice

If commercial self-hosting is mandatory, Holo4 35B-A3B is the practical candidate. Its weights use Apache 2.0; 27B uses CC BY-NC 4.0 and is licensed for non-commercial use. That can outweigh a benchmark gap when data must remain inside your network.

What Changes Between the Two Holo4 Models?

Dense 27B vs 3B-Active MoE

Holo4 27B is a dense 27-billion-parameter model. Holo4 35B-A3B is a mixture-of-experts model with 35 billion total parameters and 3 billion active per token. Both list a 262,144-token maximum context and target graphical interfaces, code, and tool calls.

Three billion active parameters can reduce compute, not storage. Capacity planning still needs the weights, image encoder, KV cache, and target context.

Hosted Access and Weight Licenses

H Models API serves holo4-27b and holo4-35b-a3b. On October 5, 2026, rates per million tokens were $0.40 input/$3 output for 27B and $0.30/$2 for 35B-A3B; cached input was $0.04/$0.03. Recheck prices before deployment.

The 27B model card labels its weights CC BY-NC 4.0. The 35B-A3B model card uses Apache 2.0. Hosted API rights remain subject to H Company’s service terms.

Compare the Decision Factors

Long-Task Reliability

This cannot be judged by feel. It needs a sample run. Build 20–50 tasks with fixed state, permissions, time limit, and verifier. Separate full completion, partial completion, wrong action, timeout, and human rescue. Keep memory, screenshots, and retry policy identical.

Cost per Attempt and Cost per Successful Task

Track two numbers. Cost per attempt includes model and infrastructure charges. Cost per successful task adds retries and operator time, then divides by accepted completions.

Using H Company’s OSWorld 2.0 figures, $1.22 divided by 41.5% implies about $2.94 per successful 27B task; $0.61 divided by 12.3% implies $4.96 for 35B-A3B. That is my arithmetic from vendor data, not a quote. Your success rate determines the answer.

Commercial Self-Hosting Rights

For commercial self-hosting, 35B-A3B’s Apache 2.0 release creates a route that the non-commercial 27B weights do not. Review attribution, notices, upstream components, and your use case with counsel. The WaveSpeed Commercial Use Policy governs that platform; it cannot replace or expand an upstream model license. This article is general information, not legal advice.

Limits and Agent Guardrails

Vendor Benchmarks Depend on the Harness

H Company reported the launch numbers in its own harness and disclosed that comparisons may use different harnesses, effort levels, and subsets. Treat them as selection evidence. Pin the revision, prompt, image settings, tool schema, step budget, timeout, and verifier before rerunning your acceptance suite.

Sandbox Tools and Require Human Approval

Give either model the smallest tool and credential scope that can complete the job. Run code in a disposable sandbox, isolate browser profiles, allowlist destinations, cap steps and spend, and log every tool action. Require human approval before payments, publishing, deleting data, changing access, sending external messages, or making other high-impact changes. Apply the WaveSpeed Acceptable Use Policy when that platform is involved, alongside H Company’s rules and your own controls.

FAQ

Does Holo4 require H Company’s own agent harness?

No. H publishes a hai-agents reference harness, but the weights can be served through vLLM or llama.cpp. You still need an agent loop to send screenshots and tool results, execute allowed actions, manage memory, and enforce approvals.

Can Holo4 operate on Android as well as desktop and web?

Yes. H Company explicitly describes the same Holo4 model operating across Android, desktop, and web. Control depends on your harness, connection, and permissions.

Which quantization formats are officially published for Holo4?

H Company publishes both sizes in BF16, FP8, NVFP4, and Q4 GGUF. NVFP4 targets compatible Blackwell hardware, while Q4 GGUF is documented for llama.cpp; validate quality after quantization.

Can Holo4 process documents as well as screenshots?

Yes, with a boundary: H’s document workflow renders each page to an image and asks Holo4 to return Markdown. The documentation says Holo reads images, not raw PDFs, so multi-page files must be rasterized and processed page by page.

Does H Company provide a public changelog for Holo4 checkpoints?

Not a dedicated, versioned Holo4 checkpoint changelog that I could verify. H publishes a platform changelog for its Agents API, SDKs, and CLI, while checkpoint changes appear through Hugging Face repository history. For a production Holo4 model selection, pin an exact revision and revalidate before upgrading.


Previous posts:

Share