ax@ax-radar:~/podcasts/bestpartners-yt $ ls -t podcasts/
33 srcsignal 72%cycle 04:32

podcasts

14 episodes · updated 3m ago
6 channels tracked
tierfeaturedallincludes low-score
最佳拍档 (BestPartners)14 episodes
2026-04-17 · Fri
09:00
158d ago
最佳拍档 (BestPartners)· atomZH09:00 · 04·17
How Hermes Agent differs from OpenClaw: Nous Research, control loop, self-improvement, and plagiarism dispute
Hermes Agent uses the agent’s own execution loop as the core, contrasting OpenClaw’s Gateway-centered design with a 4-layer memory stack and cron checks every 60 seconds. The video says Hermes keeps about 1,300 tokens of persistent memory, stores history in SQLite plus FTS5, saves skills in ~/.hermes/skills/, and supports migration from ~/.openclaw. The key shift is procedural memory, but the EvoMap plagiarism dispute is only described by the video; the post does not disclose verifiable evidence.
#Agent#Memory#Tools#Nous Research
editor take
Hermes Agent's real shift is procedural memory over facts, but the plagiarism claim is video-only with no verifiable evidence.
sharp
Hermes Agent shifts control to the agent’s own execution loop, then backs that choice with ~1,300 tokens of persistent memory, SQLite plus FTS5 history retrieval, 60-second cron polling, and skills stored as durable artifacts. I buy that direction. It targets the actual bottleneck in personal agents: factual memory has been easy for a while; procedural memory has not. Plenty of systems remember that you prefer zsh or daily briefings. Very few reliably turn a successful multi-step task into something reusable on the next run. The video frames Hermes versus OpenClaw as a split in design philosophy, and that feels broadly right. OpenClaw’s Gateway-centered architecture is strong on auditability, control, and clear workspace boundaries. Hermes puts the execution loop at the center and lets the rest of the stack orbit it. The payoff is a cleaner learning loop: complete a task, then formalize it as a skill, then reuse it later. The part I care about is not the “self-improving” slogan. It’s that skills are treated as a fourth memory layer, stored in ~/.hermes/skills/ and managed by tools inside the system. For builders, that matters more than “long-term user preferences.” Preference memory changes tone. Procedural memory changes cost structure. I’ve thought for a while that a lot of 2025-era agent products overstated what “memory” meant. They glued together RAG, logs, markdown files, and some summaries, then called it long-term learning. Hermes at least sounds structurally more serious. A tiny core memory budget of about 1,300 tokens forces prioritization. Session history in SQLite plus FTS5 signals that most context should stay off-prompt until needed. Skills as a separate layer acknowledges that “what the agent knows” and “what the agent knows how to do” are different assets. That decomposition lines up with the better research-oriented agent work. MemGPT and related systems were already wrestling with context overflow, but most implementations stopped at retrieval and summarization. Hermes tries to go one step further by turning experience into executable assets. That said, I don’t buy the stronger “self-improving” claim from the video without more evidence. Automatic skill generation is not the same as automatic improvement. If the abstraction boundary is wrong, the agent just hardens one accidental success into a brittle routine and then repeats it. Anyone who has built shell-heavy agents has seen this: the workflow works once, then the directory layout changes, a permission flag changes, an API field changes, and yesterday’s “learning” becomes today’s failure mode. The article gives no numbers on skill-generation success rate, rollback behavior, pruning rules, or reuse hit rate across long-running tasks. Without those, “gets better over time” is still a design goal, not a demonstrated system property. I also want to push back on the implicit narrative that OpenClaw’s centralized Gateway is somehow a legacy choice while Hermes’s loop-centered architecture is inherently superior. Centralization is often the price of operational sanity. Once scheduling, memory refresh, skill generation, and cron execution all sit close to the agent loop, self-reference complexity rises fast. Debugging gets uglier too. A bug in a tool call is annoying. A bug that produces a bad skill and then gets reused across future sessions is worse. The video lists five layers of security, SSRF defenses, dangerous-command prechecks, and isolation. Good. But the body still does not disclose the default permission model, the exact isolation boundary, or how credentials are handled when connected to Telegram, Discord, Slack, or WhatsApp. In self-hosted agents, security is not about how many protections you can name. It’s about whether the system defaults to denial in the places that matter. The wider context helps here. After Anthropic pushed computer-use style workflows into the mainstream, a lot of the market focused on “the model can click buttons and call tools.” That was never the hard part for sustained adoption. The hard part was whether the system developed reusable organizational memory after ten or fifty runs. OpenDevin, OpenHands, and the whole ecosystem around coding agents kept hitting the same wall: short tasks looked great; long-horizon maintenance degraded. Hermes’s layered memory plus skill accumulation is a direct answer to that wall. I haven’t personally run Hermes on a long-duration setup, so I’m not treating this as proven. But at the architecture level, it’s more convincing than just throwing a larger context window at the problem. Bigger context does not magically produce method. On the EvoMap plagiarism dispute, I’m not willing to take a position from this material alone. The title and video narration mention it, but the body does not provide verifiable evidence, commit history, or a timeline. Open-source agent projects are converging on similar directory layouts, prompt conventions, and memory patterns anyway. If you want to make a plagiarism case here, you need repository history and design chronology, not vibes. My take is simple: Hermes matters because it tries to change the unit of value in a personal agent from chat history to executable workflow memory. If that works in practice, the moat stops being “which model API do you support” and starts becoming “which system can distill failures and successes into stable reusable actions.” The video gives enough architecture to take the bet seriously. It does not yet give enough longitudinal evidence to declare the bet won.
HKR breakdown
hook knowledge resonance
open source
71
SCORE
H1·K1·R1
2026-04-16 · Thu
10:03
159d ago
最佳拍档 (BestPartners)· atomZH10:03 · 04·16
Who Is Satoshi Nakamoto? A New York Times investigation points to Adam Back, with community pushback
The video says a New York Times investigation published on April 8, 2026 points to Adam Back as Satoshi Nakamoto, based on stylometry, technical lineage, timeline gaps, and disputed emails. It cites a filter from 34,000 mailing-list users to 1 candidate, 521 overlapping terms, and 67 shared hyphenation errors; but it also says there is no genesis-key proof, Adam Back denies it, and critics dispute the code style and motive. The key point: this is an indirect evidence chain, not a cryptographic confirmation.
#Adam Back#The New York Times#John Carreyrou#Commentary
editor take
If the Times really pinned Adam Back this hard, it's still a strong suspicion case, not an identification. No key signature, no closure.
sharp
The video says the New York Times tied Satoshi to Adam Back on April 8, but it still lacks genesis-key proof. My read is simple: this sounds like an elite circumstantial case, not a technical identification. In crypto, those are different leagues. The recap throws out sticky numbers. It says 34,000 mailing-list users were filtered to one candidate. It cites 521 overlapping terms and 67 shared hyphenation errors. But I could not verify the original Times piece here, and the video does not fully disclose the methodology. What was the control set. How were false positives handled. Were the samples time-normalized. Stylometry can raise confidence. It does not survive adversarial disguise, ghostwriting, or contaminated archives on its own. There is also a strong historical reason to push back. Newsweek pointed at Dorian Nakamoto in 2014 and face-planted. HBO's 2024 film pushed Peter Todd and got heavy criticism from the Bitcoin crowd. Every Satoshi hunt eventually runs into the same wall: without a cryptographic signature from an early known key, the story remains inference. The field already settled that standard years ago. Now, Adam Back is not a random suspect. Hashcash is the clearest technical ancestor to Bitcoin's proof-of-work. Back was early enough in the cypherpunk circles. The ideological overlap also tracks. I have long thought he fits the profile better than many media-friendly suspects. But “plausible architect” is not “confirmed author.” If the article did not publish raw email metadata, reproducible archive references, and enough material for outside researchers to rerun the chain, readers are still being asked to trust a newsroom, not verify a claim. I am especially skeptical of the “his reaction was odd” angle. Great investigative reporters use behavior cues. That still ages badly in technical identity cases. Craig Wright was not discredited because his affect looked wrong. He was discredited because the cryptographic evidence collapsed. Same rule here. Sign an early message. Move an early UTXO. Produce clean, continuous provenance. Without that, even a 10,000-word investigation stays in the category of very serious suspicion.
HKR breakdown
hook knowledge resonance
open source
18
SCORE
H1·K0·R0
2026-04-15 · Wed
2026-04-14 · Tue
2026-04-13 · Mon
04:53
162d ago
最佳拍档 (BestPartners)· atomZH04:53 · 04·13
2026-04-13 livestream: Can prolonged AI use cause physiological discomfort?
This 2026-04-13 livestream centers on whether prolonged AI use causes physiological discomfort, and only the title is disclosed. The RSS snippet is empty; the post does not disclose speakers, sample size, symptom definitions, measurement methods, or conclusions.
#Commentary
editor take
Only a title — no speakers, sample, symptom definition, or conclusion. Don't let the headline lead you.
sharp
This livestream discloses 1 title and no body details on sample size, symptom definition, measurement method, or control condition. My read is simple: without those basics, any claim that “AI use causes physiological discomfort” has not cleared the evidence bar. Look, this topic invites category errors. Staring at a screen for two hours can cause eye strain. Continuous typing can cause neck and shoulder tension. Open-ended chat systems can extend session length. High cognitive load can trigger headaches or nausea. All of those are real, but they are not the same mechanism. If someone wants to show an AI-specific effect, they need a control design: same 60–90 minute task, compared across search, document editing, coding IDEs, and a chat model, while holding screen brightness, break frequency, typing volume, and task difficulty roughly constant. The title gives none of that. There is also useful context outside the article. Over the past year, we have seen headlines around “ChatGPT psychosis,” emotional dependency on chatbots, and AI-induced distress. The claims that held up were usually case reports, clinical cautions, or survey correlations. They were not clean physiological mechanism studies. In adjacent HCI areas, reproducible findings on screen fatigue, notification load, or VR sickness usually come with clear operational definitions and experimental setups. A title alone does not earn that credibility. My pushback is that this framing can hide a product problem inside a media panic. If the discomfort comes from latency in voice mode, hallucinations creating cognitive dissonance, or compulsive interaction loops from chat UX, then the target is the interaction design, not “AI” as a single causal object. Right now, only the title is disclosed. I can accept this as a valid research question. I do not accept it as an established finding.
HKR breakdown
hook knowledge resonance
open source
24
SCORE
H1·K0·R0
2026-04-12 · Sun
23:00
163d ago
最佳拍档 (BestPartners)· atomZH23:00 · 04·12
Sam Altman's Many Faces: New Yorker report, internal documents, and the OpenAI firing saga
This YouTube video says The New Yorker spent 18 months, interviewed 100+ people, and cited two internal documents to examine Sam Altman and OpenAI governance disputes. The post also mixes in unresolved lawsuits and allegations; it does not provide independently verifiable source materials, so the key watchpoints are board failure, Microsoft tensions, and Superalignment resource allocation.
#Alignment#Safety#Sam Altman#OpenAI
editor take
The New Yorker's 18-month investigation paints Sam Altman as a serial liar who gutted OpenAI's safety promises for power and profit.
sharp
The claimed fact pattern here is large: The New Yorker reportedly spent 18 months, interviewed 100+ people, and relied on 2 internal documents. If that sourcing holds up, this is not celebrity gossip. It is another stress test showing that OpenAI’s original promise — nonprofit governance restraining commercial acceleration — largely stopped working by late 2023. The video spends a lot of energy on Sam Altman’s character, alleged lying, old YC stories, and personal drama. I don’t think that is the core read. The core read is structural: a board removed a CEO in November 2023, failed to hold the line for even 5 days, and then accepted a settlement that left the CEO stronger than before. That is what institutional failure looks like. The sharpest operational claim in the video is the Superalignment gap: public messaging around 20% of compute, internal reality allegedly at 1% to 2%. That number matters because we already had a strong public breadcrumb. Jan Leike said in 2024, under his own name, that safety culture and processes had taken a back seat to “shiny products.” That was not an anonymous whisper. So the broad direction here matches what the field already suspected. OpenAI’s 2024–2025 cadence was product first: enterprise features, multimodal rollout, voice, API monetization, deeper distribution. A safety team getting squeezed is not surprising under that pressure. The issue is the mismatch between the institution’s self-description and its budget allocation. If the brand says “safety-first lab” and the compute ratio lands closer to 2% than 20%, outsiders should treat the safety story as recruiting and legitimacy infrastructure unless the company shows receipts. I also have pushback on the video itself. It mixes unresolved litigation, assault allegations, old interpersonal accounts, Microsoft tensions, and New Yorker reporting into one continuous moral narrative. That is exactly where careful source separation matters, and the post does not provide a source pack for the two documents it says exist. No raw memo, no notes appendix, no clean boundary between magazine reporting, court filings, public tweets, and the channel’s own interpretation. That makes a big difference. Since the November 2023 board crisis, the Sam narrative has split into two camps: one says he is the only executive who can turn frontier research into products at global scale; the other says he is a power center governance cannot constrain. Both camps have evidence. Without primary materials, I’m not signing off on a full conviction narrative from a YouTube retelling. There’s also a wider context the video only partially captures: OpenAI’s problem was never just Sam, and it was never just a weak board. The hybrid structure was unstable from the start. A nonprofit parent claimed a mission to humanity, while the operating engine depended on massive commercial capital and Microsoft cloud support. That arrangement could survive when the company was still a research lab. After GPT-4 and the revenue explosion, it needed unusually strong information rights, escalation rules, and investor firewalls. I haven’t seen evidence that those controls were ever built well enough. Once that’s true, any CEO with product traction, employee loyalty, and investor backing will overpower the board. Anthropic is the obvious comparison. I’m not romanticizing it; every frontier lab eventually faces the same compute-and-revenue gravity. But Anthropic’s pitch has at least stayed more coherent around safety process, external policy engagement, and capital raised explicitly for frontier training. OpenAI tried to preserve a mission-governed identity while becoming the market’s most important consumer AI company. That tension was always going to snap somewhere. So my take is not “Sam is good” or “Sam is evil.” That frame is too easy. The harder question is who controls the compute budget, who can override safety allocation, and who survives when the board, investors, employees, and strategic partner all pull in different directions. If the answer keeps being “the CEO,” then OpenAI’s long-running governance story has been far thinner than its public positioning.
HKR breakdown
hook knowledge resonance
open source
39
SCORE
H1·K0·R1
09:01
163d ago
最佳拍档 (BestPartners)· atomZH09:01 · 04·12
Buffett's first CNBC interview after stepping down: charity lunch auction returns; Abel, Apple, Fed and nuclear risk
Buffett said in his first CNBC interview after stepping down as Berkshire CEO that he will restart the charity lunch auction halted in 2022. The auction runs from May 7 19:30 to May 14 19:30 PT; it raised over $50 million across 22 years, with the last sale at $19.1 million, and Buffett said Berkshire still holds over $350 billion in cash and Treasuries, including $17 billion bought that week. The sharper signal is his pricing discipline: he said the market pullback is still not attractive, and Berkshire has made over $100 billion on Apple.
#Warren Buffett#Berkshire Hathaway#CNBC#Commentary
editor take
Buffett's first post-CEO interview restarts the charity lunch auction, but the real signal is his pricing discipline: the pullback isn't cheap enough.
sharp
Buffett kept more than $350 billion in cash and Treasuries, and he bought another $17 billion that week. My read is simple: that matters far more than the revived charity lunch. A 95-year-old allocator with unlimited patience still does not like current prices after a pullback. For anyone working around AI, that is the signal. The market spent the last year treating AI capex, model demand, and platform concentration as enough to justify almost any multiple. Buffett is saying no with actual balance-sheet behavior. I think AI markets keep blurring two separate claims. One claim is that AI demand is real. That looks true. The other claim is that current public-market prices still offer good odds. Buffett is attacking the second claim. The article gives two hard anchors: Berkshire still sits on $350 billion-plus in cash and Treasuries, and Apple has already made Berkshire more than $100 billion. Put together, those numbers say he is not anti-tech and not afraid of size. He just refuses to add when expected return no longer clears his hurdle. That matters because a lot of AI positioning over the last year has been sold as conviction when it often looked like momentum with a thesis attached. Buffett's Apple comments reinforce that. He said he does not regret trimming because the position had become too large relative to the rest of the portfolio. That is portfolio discipline, not a macro call. Many AI-heavy funds did the opposite. They let one theme dominate and then reframed concentration as expertise. Sometimes that is skill. Sometimes it is just what happens when one trade keeps going up. I have one pushback on Buffett's line that he will not play AI because he does not understand it and is late. As a personal rule, fair enough. As a description of economic exposure, it is incomplete. Berkshire already has indirect AI exposure through Apple, through the rate environment that rewards short-duration Treasury holdings, and through the broader concentration of US equity returns in tech-heavy giants. So this is not an abstention from the AI era. It is a refusal to underwrite technology-path risk directly. He is choosing cash yield and proven cash flows over frontier uncertainty. The outside context makes this sharper. Over the last year, Microsoft, Meta, Alphabet, and Amazon all kept lifting capex. By memory, their combined annualized spend is in the several-hundred-billion-dollar range now, though I have not rechecked the latest filings. Public markets have largely accepted that spending because investors assume AI revenue and margins will catch up later. Buffett's posture is a reminder that demand can be real while equity still gets overpriced. We learned that lesson in earlier cycles. The internet was real in 2000. Plenty of stocks were still too expensive. AI today is sturdier than that era in many ways. Revenue quality is better. Deployment is broader. But the distinction between a good business and a good entry price still holds. I also do not fully buy the interview packaging. The headline crams in philanthropy, inflation, nuclear risk, Gates, Epstein, and succession. That is good television. It is not the core signal. The useful unanswered questions are elsewhere: what maturities Berkshire is buying in Treasuries, how much investment authority Abel now has in practice, and whether Buffett has an explicit valuation framework for the other megacap platforms beyond Apple. The body does not disclose those details, so I will not pretend it does. With the facts we do have, the takeaway is blunt: if Buffett still finds this drawdown uninteresting, the market is still paying up for certainty that has not fully been stress-tested.
HKR breakdown
hook knowledge resonance
open source
6
SCORE
H0·K0·R0
2026-04-11 · Sat
09:00
164d ago
最佳拍档 (BestPartners)· atomZH09:00 · 04·11
AI Is Accelerating: Greg Brockman on 70% AGI, Spud, Sora, and the Super App
According to the video’s retelling, Greg Brockman said OpenAI sees the path to AGI as 70% to 80% complete, and the new pretrained base model Spud has finished pretraining. The post also says OpenAI is pausing broad Sora expansion because of compute limits and is prioritizing GPT reasoning models, a super app, and an automated AI researcher targeted for this fall; it frames a $110B infrastructure buildout as a revenue center. The post does not disclose the original interview date, Spud specs, benchmark results, or release timing.
#Reasoning#Code#Agent#OpenAI
editor take
Greg Brockman claims AGI is 70–80% done and new base model Spud finished pretraining, but the post doesn't disclose specs, benchmarks, or release timing.
sharp
OpenAI ties a reported $110B infrastructure buildout to the GPT line, while Sora gets slowed by compute limits. My read is simple: the useful signal here is not the “70% to 80% to AGI” claim. It is the resource allocation logic. OpenAI appears to be prioritizing products that monetize fast, retain daily users, and compound usage inside one interface. I do not buy the “AGI is 70% to 80% complete” line as an external metric. The retelling gives no original interview date, no task suite, no failure boundary, and no cost threshold. The article defines AGI as human-like competence at operating computers for knowledge work. Fine. By that definition, the field has moved a lot over the last year. Anthropic pushed coding and agents, Google kept folding Gemini into tool use and multimodal workflows, and OpenAI has been turning coding ability into a broader assistant product. But turning that into a percentage is internal morale language, not a reproducible benchmark. I do find the Sora deprioritization plausible. Video generation burns training and inference compute, while user value per unit of compute is still less obvious than coding, office tasks, search-like assistance, and enterprise workflows. If OpenAI has a stronger base model in the pipeline and still needs RL, post-training, deployment, and ChatGPT capacity at scale, compute will flow to the main line first. That is not unusual. Across the last year, major labs kept moving flashy demos behind tools that fit into recurring workflows and recurring revenue. The “unified GPT architecture” claim needs pushback. The article says text, voice, and image all sit under one GPT-style core, and even image generation is framed as part of that line rather than a separate diffusion-first stack. I believe half of that. Product unification is real across the industry. Users increasingly interact with one system, not a visible bundle of models. But product unification is not the same as training unification. The body gives no architecture details, no loss design, no routing, no benchmarks, and no cost data. Without that, nobody outside the company can tell whether this is one base model or several specialized subsystems wrapped into one GPT experience. Spud is still mostly a placeholder. The article only says pretraining is done and that Spud is a new foundation model for later RL and post-training. That description is generic and believable. It also tells us almost nothing. No parameter scale is disclosed. No token count is disclosed. No context window, benchmark, release timing, or relation to existing model families is disclosed. So the key question stays open: is Spud a genuine generational jump, or a fresh inventory layer for products and internal distillation? The title gives a name. The body does not give a role. The “super app” part is the most credible strategic piece here. ChatGPT stopped being a pure chatbot business a while ago. The market has been teaching the same lesson for two years: users do not pay for “a bit smarter” by itself. They pay when AI removes steps, reduces tool switching, and takes ownership of workflow fragments. Anthropic pushed Claude into coding and enterprise use. Microsoft kept embedding Copilot into Office. Google keeps using Search and Workspace as distribution. If OpenAI is trying to combine memory, browsing, coding, spreadsheet work, and delegated action into one front end, that is not a novel idea. It is still the clearest path to retention and higher revenue per user. The hard part is not the model. It is permissions, reliability, rollback, auditability, and interface design. The automated AI researcher claim deserves caution. AI systems already help with literature review, experiment drafting, and result analysis. Calling that an end-to-end researcher targeted for this fall is a stronger statement. I would discount it until we see scope and evaluation. Over the last year, many “AI scientist” systems looked impressive on constrained benchmarks, then weakened on messy data, failed experiments, open-ended hypotheses, and interpretation under uncertainty. Treat it like a high-throughput research intern and the claim sounds reasonable. Treat it like an autonomous scientist and the article does not provide enough evidence. The safety section also pulls in two directions. It stresses prompt injection and alignment work, then leans on openness and resilience as governance language. I have doubts there. OpenAI’s actual product posture over the last two years has not been especially open at the frontier-weight level. “Broad participation” works as a governance value statement. It does not map cleanly onto current practice. The article provides no new evals, no red-team numbers, and no misuse interception rates, so I would not treat this as evidence of safety progress. My bottom-line read is narrow. Three things are believable: OpenAI still has severe compute scarcity, GPT remains the internal priority, and product usability has become a first-order concern. Three things should not be accepted at face value: the AGI percentage, Spud’s significance, and the automated researcher timeline. Without the original interview, benchmarks, or release details, those claims are still narrative, not proof.
HKR breakdown
hook knowledge resonance
open source
64
SCORE
H1·K0·R1
2026-04-10 · Fri

more

feeds

admin