Fable 5.1 Computer Use: What OSWorld Scores Mean
Fable 5.1 computer use scores explained for builders evaluating browser and desktop agents, with the benchmark limits kept visible.

When a launch chart gives me 77.9% and 41.7% for the same task set, I stop at the labels. That is the useful tension in this Fable 5.1 computer use review: partial progress looks strong, while strict completion is much lower.
Dora. I checked Anthropic’s launch materials. I did not reproduce the vendor run. This is a source-based reading, not a production claim.
What the OSWorld Result Measures
Anthropic’s Fable 5.1 system card reports 77.9% partial credit and a 41.7% strict pass rate on OSWorld 2.0. Both are Pass@1 averages across five independent runs. The setup used 1080p, a 500-action limit, maximum reasoning effort, and Claude Opus 4.8 as grader where a model grader was required.

The distinction matters:
| Metric | Published result | Operational reading |
|---|---|---|
| Partial score | 77.90% | Average checkpoint credit within each task |
| Strict pass rate | 41.70% | Tasks where every required checkpoint passed |
The benchmark paper defines 108 long-horizon workflows on a live Ubuntu virtual machine. The agent receives screenshots and acts through mouse and keyboard controls. This computer use benchmark measures end-to-end state changes. It does not isolate visual grounding, click precision, planning, or recovery.
Why Safeguards Affect the Score
Anthropic ran Fable 5.1 with production safeguards enabled. Certain interventions scored zero; other flagged cyber and biology tasks went to fallback models. The result describes a policy-controlled route, not an unfiltered model.
Refusals and Fallback in Computer Tasks
A refusal can be the correct safety outcome and still count as failure. A fallback may finish, but the executing model has changed. I paused here. Log the requested model, served model, safeguard category, refusal point, retries, and final state. Otherwise completion can hide which system did the work.
Task-Version Comparability
Anthropic used the benchmark authors’ August 2026 task release, later task fixes, and fixes to setup and grading scripts that Anthropic reported upstream. Fable 5 and Opus 5 were rerun under the same conditions. Earlier published results used different task files and cannot be compared directly. A score belongs to the model, task release, harness, effort level, grader, and safeguards together.
What the Result Means for Product Teams
The result is a shortlist signal for long-running GUI work, not evidence that Fable is generally best at browsers or desktops. Complete delivery remains the harder problem. Partial progress helps only when a person or recovery loop can finish the work.
Browser and Desktop Agent Fit
The evaluation used Ubuntu and screenshot-driven controls. It does not establish identical behavior on Windows, macOS, a Chrome extension, or a DOM-aware Claude browser agent. Authentication, native dialogs, extensions, and display scaling change the action loop. Reversible, reviewable background tasks are the closest fit. Financial approval and credential-heavy flows need separate evidence.
Reliability Checks Beyond One Score

Track strict completion, unintended side effects, repeated actions, recovery attempts, step count, latency, token use, refusals, fallbacks, and prompt-injection exposure. The AWS computer-use guidance also calls for an isolated low-privilege environment, restricted network access, and human oversight for sensitive actions. Those controls belong in the agent evaluation, not in a later security checklist.
Reproduce a Narrow Internal Computer-Use Test
Use one staging expense workflow. Start from a clean VM snapshot. Ask the agent to open a supplied receipt, create an expense draft, enter the vendor, date, amount, and category, attach the receipt, save the draft, and verify the audit row. It must not submit the expense.
Pin the model ID, tool version, OS image, browser build, display size, effort, and step cap. Define the strict pass first: every field correct, attachment present, audit entry present, no submission, and no unrelated state change. Run 20 trials from the same snapshot. Store actions, screenshots, served model, fallback events, tokens, wall time, and final state. Partial checkpoints explain failures. Strict completion decides whether the route advances.
FAQ
Does Anthropic publish computer-use incident reporting guidance for Fable 5.1?
I found no Fable-specific computer-use incident playbook in the launch page or system card. Teams still need an internal route covering containment, trace preservation, affected accounts, model and tool versions, provider escalation, and user notification. Security vulnerabilities can use Anthropic’s general disclosure channel; operational mistakes need the provider’s support process.
Are screenshots and action traces publicly available?
Not for Anthropic’s Fable run, based on the materials reviewed. The system card publishes aggregate scores and settings, not the five run-level screenshot and action histories. Benchmark authors publish selected task trajectories, but those are not substitutes for Anthropic’s exact evaluation artifacts. This is where my data ends.
Did the evaluation use a custom agent harness?
Anthropic does not label it a custom harness. The card specifies a live Ubuntu VM, screenshot and mouse-keyboard interaction, default resolution and step limit, maximum effort, and a named grader. It does not fully publish the system prompt, orchestration logic, or complete run artifacts. An exact reproduction is therefore not currently possible from the card alone.
Which cloud providers expose Fable 5.1 computer-use tooling?
Anthropic’s compatibility table lists Claude API, Claude Platform on AWS beta, Amazon Bedrock beta, Google Cloud, and Microsoft Foundry beta. The Google Cloud model card explicitly lists computer use for Fable 5.1. Availability still varies by tool version, endpoint, and region. Model access alone does not prove tool parity.

What data-retention rules apply to computer-use screenshots and recordings?
The client environment stores screenshots, actions, and files as tool artifacts, while screenshots sent in API requests are processed as request content. Anthropic’s API retention documentation classifies Fable 5.1 as a Covered Model requiring 30-day retention by default unless Anthropic expressly authorizes otherwise. Bedrock and Google Cloud apply provider-controlled handling, so contracts and trace-storage settings still need separate review.
Conclusion
The honest reading is narrow. Fable 5.1 computer use made substantial checkpoint progress and completed every checkpoint in 41.7% of evaluated task runs. That supports an internal trial, not a production verdict. Pin the full test configuration, score strict completion, record safeguards and fallback, then decide from the workflow that actually matters.
Previous posts:





