ax@ax-radar:~/all $ grep -v 'tier=excluded' stream.log
33 srcsignal 72%cycle 04:32

posts · 2026-08-16

27 items · updated 3m ago
RSS live
2026-08-16 · Sun
23:45
37d ago
● P1Hacker News Frontpage· rssEN23:45 · 08·16
Alibaba releases Qwen 3.8 27B model with strong performance but excessive default reasoning
Simon Willison tested Alibaba's Apache 2 licensed Qwen 3.8 27B. The 17GB quant runs on a laptop and nails SVG generation and bounding boxes, but the default xhigh reasoning setting is a trap: a simple 'draw a circle' prompt triggered minutes of overthinking and an animated geometric study. A pelican-on-a-bike SVG took 21 minutes and 22,276 reasoning tokens; with reasoning off, the same prompt finished in two minutes. Start with low or no reasoning.
#Reasoning#Vision#Code#Alibaba Qwen
why featured
Featured · importance 88 · hook + knowledge + resonance
editor take
Qwen 3.8 27B is a solid model, but the default xhigh reasoning setting is a self-inflicted wound — it'll overthink a circle for minutes. Turn reasoning off first.
sharp
Qwen 3.8 27B dropped Friday — Apache 2 licensed, 27B params, vision-capable, and sized right for a decent laptop. Both sources point to Simon Willison's hands-on testing, not an official press release, so the coverage is consistent but narrow in perspective. Willison ran the 17GB Q4_K_M quant on an M5 Max MacBook Pro and an NVIDIA DGX Spark. The model defaults to xhigh reasoning effort, and it shows: generating a pelican-on-a-bicycle SVG burned 22,276 reasoning tokens and took 21 minutes. Same prompt with reasoning off took two minutes — worse image, but functional. The real comedy was asking for a circle. The model launched into a multi-minute deliberation about color palettes and geometric aesthetics, then produced a beautiful animated circle that was absolutely not what was requested. I'd read this as a default-config faceplant, not a model capability problem. Qwen's self-reported benchmarks show gains over both the 3.6 27B and the closed-weight 3.7-Plus, but independent evals aren't in yet. What's missing: side-by-side comparisons across reasoning levels, and API pricing. If you try this model, step one is setting reasoning_effort to low or off.
HKR breakdown
hook knowledge resonance
open source
88
SCORE
H1·K1·R1
22:33
37d ago
r/LocalLLaMA· rssEN22:33 · 08·16
Simon Willison: Qwen 3.8 27B is excellent, but defaults to wildly overthinking
Simon Willison finds Qwen 3.8 27B excellent but notes it defaults to overthinking, producing overly verbose outputs. The post does not disclose specific benchmarks or baselines.
#Simon Willison#Qwen
editor take
Simon Willison says Qwen 3.8 27B is great but defaults to verbose, essay-like answers.
HKR breakdown
hook knowledge resonance
open source
60
SCORE
H1·K1·R0
20:31
37d ago
● P1Hacker News Frontpage· rssEN20:31 · 08·16
Stripe acquires AI model routing platform OpenRouter for over $7 billion
Stripe is finalizing a $7B+ deal to buy OpenRouter, an API routing layer that lets devs call 300+ models with usage-based billing. It's Stripe's largest acquisition yet, pulling the payments giant straight into AI infra. OpenRouter handled 150B model requests last year with ~200K monthly active devs. The post doesn't spell out the closing timeline or cash-vs-stock mix.
#Stripe#OpenRouter
why featured
Featured · importance 100 · hook + knowledge + resonance
editor take
Stripe's $7B+ OpenRouter deal isn't about buying models — it's about owning the checkout counter where every API call eventually settles. Smart, but the price tag needs scrutiny.
sharp
This story hit Bloomberg, FT, TechCrunch, and HN's front page simultaneously — the coverage density alone says the market is taking it seriously. OpenRouter is a model router: you send a request, it picks the best model for your budget and latency needs. Sounds like middleware, but the real asset is traffic data — who's calling which model, at what cost, with what latency. Stripe buying it means they're planting a flag at the payment layer of every AI application. The price numbers don't fully line up: Bloomberg says over $7B, FT says $8B, TechCrunch says "$7B+ reportedly." That gap isn't trivial, but all sources trace back to the same anonymous tipsters — no official announcement yet. I'd discount the exact figure until Stripe confirms. One TechCrunch piece ran the headline "Stripe didn't really buy OpenRouter because of the 'singularity'" — a useful pushback against reading this as AI hype. It's an infrastructure play. What's missing: OpenRouter's revenue, margins, and how Stripe plans to integrate it. If the platform moves huge volume but takes a thin cut, the $7B math needs a different lens.
HKR breakdown
hook knowledge resonance
open source
100
SCORE
H1·K1·R1
18:48
37d ago
Hacker News Frontpage· rssEN18:48 · 08·16
Protobuf finally has LSP support with Buf's production-grade server
Buf released the first production-grade LSP server for Protobuf, bundled in the Buf CLI. It uses a new query-driven compiler frontend for incremental compilation and precise diagnostics, like catching duplicate modifiers. Works with VSCode and Neovim. Future features include auto-imports, custom option completion, and field number suggestions.
#Buf#VSCode#Neovim
editor take
Buf shipped a production-grade LSP for Protobuf, works with VSCode and Neovim out of the box.
HKR breakdown
hook knowledge resonance
open source
55
SCORE
H0·K1·R0
16:53
37d ago
TechCrunch AI· rssEN16:53 · 08·16
Anthropic CEO says AI backlash is 'fundamentally a crisis of trust'
Anthropic CEO Dario Amodei pushes back against claims that his warnings fueled the AI backlash. Investor Gavin Baker argued Amodei's risk narrative hurt public trust in AI infrastructure. Amodei calls the backlash a crisis of trust, not technology. The post doesn't spell out his full counterargument.
#Anthropic#Dario Amodei#Gavin Baker#Policy
editor take
Amodei calls AI backlash a trust crisis, not a tech problem—but the post doesn't spell out his full counterargument.
HKR breakdown
hook knowledge resonance
open source
62
SCORE
H1·K0·R1
14:56
37d ago
The Verge · AI· rssEN14:56 · 08·16
ChatGPT’s Computer History tracks your clicks and keystrokes
OpenAI added a 'Computer History' feature to ChatGPT that logs clicks and keystrokes, similar to Microsoft Recall but without screenshots. The author calls it still creepy, just less so. The post doesn't disclose whether users can disable or delete the logs.
#OpenAI#Microsoft
editor take
ChatGPT now logs your clicks and keystrokes like a less creepy Recall. The post doesn't say if you can turn it off.
HKR breakdown
hook knowledge resonance
open source
55
SCORE
H1·K0·R1
14:44
37d ago
Hacker News Frontpage· rssEN14:44 · 08·16
The AI Credit Resale Economy: Who Are the Token Brokers
Vectoral investigated the token broker market where unused startup API credits are resold at 40–80% off list price. One broker claimed $100k daily spend capacity, routing requests through a proxy instead of sharing raw keys. Public marketplaces like AI Credits and CheapCredits have emerged, alongside Telegram and Reddit channels. The author estimates tens of millions of dollars in credits are circulating and expects provider crackdowns as cost awareness grows.
#Vectoral#AI Credits#AICreditMart
editor take
A secondary market for unused startup API credits is emerging, with brokers reselling at 40–80% off and one claiming $100k daily spend capacity.
HKR breakdown
hook knowledge resonance
open source
72
SCORE
H1·K1·R0
13:21
37d ago
Hacker News Frontpage· rssEN13:21 · 08·16
A public AI whose memory is shared across all users
Wild Static lets everyone talk to the same AI with no memory isolation or reset—later visitors see what was left before. The page warns “you're gonna ruin it” and tells people not to say anything private. The post doesn't disclose the underlying model, memory cap, or moderation approach; it reads more like a social experiment than a product.
#Plural Matter
editor take
Everyone shares one AI memory—the page warns you'll ruin it. Don't use it as a tool or share anything private.
HKR breakdown
hook knowledge resonance
open source
48
SCORE
H1·K0·R1
10:00
37d ago
Financial Times · Technology· rssEN10:00 · 08·16
Big Tech's data center boom set to drive up carbon emissions
The FT reports that Big Tech's massive data center buildout will increase carbon emissions. AI training and inference demand is driving a surge in electricity consumption, while clean energy deployment can't keep pace. The post doesn't disclose specific emission figures or company targets, but notes the conflict with earlier net-zero pledges.
#Financial Times
editor take
FT: Big Tech's data center power surge is outpacing clean energy, pushing emissions up.
HKR breakdown
hook knowledge resonance
open source
55
SCORE
H0·K0·R0
03:22
37d ago
Hacker News Frontpage· rssEN03:22 · 08·16
ProofRun: a local verification receipt for AI coding agents
ProofRun is an open-source tool that lets AI coding agents produce a locally signed execution receipt. It bundles the agent's environment, commands, output, and exit code into a JSON receipt and signs it with Ed25519 to prevent later repudiation or tampering. macOS and Linux are supported; Windows is not yet implemented. The project is early-stage—the post doesn't disclose performance overhead or real-world deployments, so treat it as a prototype for now.
#Agent#yebiguo#ProofRun
editor take
ProofRun signs each agent command execution with Ed25519 into a JSON receipt — tamper-proof audit trail for AI coding agents.
HKR breakdown
hook knowledge resonance
open source
62
SCORE
H1·K1·R0
00:10
38d ago
Hacker News Frontpage· rssEN00:10 · 08·16
Big Pickle scores 50.8% on SWE Atlas Codebase QnA, open-sources full configs and logs
A model codenamed Big Pickle (OpenCode Zen stealth model) scored 50.8% on Scale AI's SWE Atlas Codebase QnA benchmark. Author PhillipChaffee released full results, verifier logs, and reproduction configs on GitHub. The post does not disclose model size, training data, or inference cost—only that it's a stealth model from OpenCode Zen.
#Code#Scale AI#OpenCode Zen#PhillipChaffee
editor take
A stealth model called Big Pickle hits 50.8% on SWE Atlas Codebase QnA, but no model size or cost disclosed—keep expectations in check.
HKR breakdown
hook knowledge resonance
open source
55
SCORE
H0·K1·R0
00:00
38d ago
● P1Computing Life · Share (鸭哥 research reports)· rssZH00:00 · 08·16
Automated Claude maintenance routine opens 388 pull requests over weeks
Boris Cherny's team ran a daily Claude routine that opened 388 PRs over several weeks, with 180 merged into main. The key isn't model smarts—it's the trigger, acceptance criteria, and review funnel working together. Cherny moved the trigger out of chat windows and into a cron job; when output missed the mark, they adjusted the routine definition instead of patching code. The post doesn't disclose whether the 208 unmerged PRs were rejected, duplicated, expired, or queued. A 46.4% merge rate shows candidate submissions naturally outpace actual merges—the review funnel is part of the design.
#Agent#Boris Cherny#Anthropic#Claude
why featured
Featured · importance 88 · hook + knowledge + resonance
editor take
388 PRs isn't a model story—it's an infra story. A cron job plus acceptance criteria let Claude run its own maintenance pipeline, and the 46.4% merge rate shows the review funnel is part of the des...
sharp
Boris Cherny ran an experiment at Anthropic: a dedicated Slack channel where Claude runs daily maintenance routines across iOS, Android, desktop, web, CLI, and Agent SDK. Over a few weeks, it submitted 388 PRs, with 180 merged into main after both automated and human review. Both sources covering this agree on the framing—it's about the mechanism, not the model—which suggests the narrative comes directly from Cherny's public posts rather than independent media interpretation. I'd unpack that 46.4% merge rate carefully. The denominator is submitted PRs, not total problems found or total agent runs. The 208 unmerged PRs could be rejected, duplicated, stale, or still in queue—the public material doesn't say. This number isn't an accuracy metric, but it does confirm something useful: candidate submissions naturally outnumber actual merges, and the review funnel is built into the system from the start. What's actually portable here isn't Anthropic's internal architecture—they haven't disclosed it. What you can borrow is the mounting pattern: pick a tedious task with clear acceptance criteria that was previously too hard to hard-code, attach a cron trigger or webhook, and write the completion conditions in a format the agent can read. Permissions matter more than prompts: give it PR-submit access only, no merge rights, and definitely no production credentials. The Clinejection and PocketOS incidents both blew up because of permission surfaces, not prompt quality.
HKR breakdown
hook knowledge resonance
open source
88
SCORE
H1·K1·R1

more

feeds

admin