FEATUREDAI HOT (Curated Pool)· aihot-apiZH14:00 · 08·10
→a16z answers with data: Can agents really use a computer yet?
a16z's Fabrizio Serafini, Seema Amble, and Eric Zhou track computer-use agents on the OSWorld-Verified benchmark. A year ago the best model scored ~30%; now Claude Fable 5 hits 85%, above the human baseline of 72%. The post argues the model is no longer the main bottleneck—the frontier is shifting from 'can the agent use a computer?' to 'can it reliably do this job inside a real company,' covering permissions, process knowledge, error handling, and caching. Production deployments exist for standardized back-office work, but agents still break when tasks drift off the runbook and costs don't work everywhere.
#Agent#Benchmarking#a16z#Fabrizio Serafini
why featured
Featured · importance 82 · hook + knowledge + resonance
editor take
Claude Fable 5 hits 85% on OSWorld-Verified, above the 72% human baseline—a16z says the model bottleneck is gone, now it's an engineering fight.
sharp
This piece is worth opening because a16z plots a steep climb on OSWorld-Verified: ~30% a year ago, now Claude Fable 5 at 85%, above the 72% human baseline. Their call is blunt—the model is no longer the bottleneck, and the fight has moved from "can the agent use a computer?" to "can it reliably do this job inside a real company?"
I'd discount the score a bit. OSWorld tests standardized desktop tasks; real back-offices have messy permissions, missing process docs, and edge cases that break agents fast. The post admits agents get brittle when work drifts off the runbook, and costs don't pencil out everywhere when caching is intractable.
The useful bit is where they see production deployments happening: standardized back-office work—updating legacy systems, moving data through portals, processing tickets—where no clean API exists and manual clicking is the baseline. If you're tracking computer-use agent timelines, the framework matters more than the score. The question isn't whether the model can click the right button; it's whether engineering teams can stitch together permissions, process knowledge, error handling, and caching into something that doesn't fall over.
HKR breakdown
hook ✓knowledge ✓resonance ✓