FEATUREDComputing Life · Share (鸭哥 research reports)· rssZH00:00 · 08·23
→GLM-5.3 tops open-source chart, Claude watermark, Anthropic's 4.4x cost, OpenAI disbands safety team
Four AI stories this week lost key details in transmission. GLM-5.3 scored 60 on Artificial Analysis's Intelligence Index, tying Kimi K3 for first among open-source models, but open weights are delayed to around Aug 28 after the team found emergent exploit capabilities. Anthropic rolled out text watermarking globally for Claude; the mark is live but no detection API exists yet, so removal tools can't prove they work. Vercel's report shows Anthropic's average token price is 4.4x other labs—not because same-tier models cost more, but because Anthropic has no ultra-cheap entry model, concentrating all volume in mid-to-high tiers. OpenAI disbanded its Preparedness team in late July, the third independent safety team dissolved in two years; FT broke the story and OpenAI hasn't publicly addressed the details.
#Z.ai#Anthropic#OpenAI
why featured
Featured · importance 78 · hook + knowledge + resonance
editor take
Four stories with the missing qualifiers restored: GLM-5.3's weights are delayed, not its API; Claude's watermark is live but detection isn't; the 4.4x premium comes from having no cheap model; Ope...
sharp
These four stories got pretty mangled in transmission this week. Here's what the original sources actually say.
GLM-5.3 scored 60 on Artificial Analysis's Intelligence Index, tying Kimi K3 for the open-source lead. The API went live on August 14, no delay there. What's delayed until around August 28 are the open weights, and the reason is specific: during post-training, the team found the model had unexpectedly learned multi-step exploit chains, with ExploitBench scores jumping from 24.4% to 54.4%. Z.ai decided to do safety hardening first. TechTimes noted this is the first time a Chinese frontier lab delayed a release due to emergent capabilities rather than export controls or compliance. The gap between open-source and closed-source leaders is down to 3 points. Both THE DECODER and Franklin Templeton are making the same call: the model is becoming a commodity, the system is the moat.
Anthropic's text watermark went live globally on August 11, but the detection API is still described in future tense: "we will soon be offering." This creates a weird situation where removal tools popped up everywhere on GitHub, but since no external detector exists, none of them can prove they work. BleepingComputer confirmed this. John Gruber called it "text adulteration" on Daring Fireball. Princeton's Narayanan flagged the real issue: zero transparency on who gets access to the verifier.
Vercel's 4.4x number comes from actual July routing traffic on their AI Gateway: Anthropic took 65% of spending on 30% of token volume. A lot of people read this as Claude being 4.4x more expensive at the same tier, but CloudZero's pricing comparison shows Opus 5 is actually 11% cheaper than GPT-5.6 Sol, and Sonnet 5 is roughly even with Terra and Gemini 3.1 Pro. The gap is entirely at the bottom: Haiku costs $2, while OpenAI's Luna is $0.45 and DeepSeek is $0.32. Vercel's own line: "Anthropic has no model at the bottom of the market." For developers, the real cost lever isn't switching providers—it's routing downgradable tasks to lightweight models.
OpenAI disbanded its Preparedness team in late July, per an FT exclusive. This team assessed catastrophic risks across bio, chemical, nuclear, and cyber domains. Responsibilities were split into existing product teams, with OpenAI framing it as pre-IPO organizational streamlining. heise and TNW both confirmed this is the third independent safety team OpenAI has dissolved in two years. OpenAI hasn't publicly addressed the specifics of the FT report.
HKR breakdown
hook ✓knowledge ✓resonance ✓