OpenAI says it will be public by 2027, Anthropic may file this month, and the private frontier-lab era suddenly looks like it is packing a suit for Wall Street.
Wall Street used to be where AI labs went looking for money. Now it is starting to look like the finish line. Beyond the IPO race teed up below, Nvidia found a very 2026 way to spend $7 billion around an AI startup without technically buying it, Stripe bought the switchboard that routes developers across hundreds of models, and data centers became enough of an election liability that Senate Republicans privately warned the industry. Meanwhile, robots learned new physical tasks from seconds of demonstration, and one independent builder trained a billion-parameter Kimi-style model for about the price of a nice dinner for four in San Francisco. Nothing says AI adolescence like an IPO, a zoning fight, and a $252 language model. Let’s get into it.
Around the Horn — Thursday, August 20, 2026
The big news today was that the two biggest private frontier labs are starting to put dates on Wall Street. OpenAI CFO Sarah Friar told employees the company would be public in 2027, or sooner if growth keeps up; Andrew Curran’s recap highlighted a 35% quarter-to-date revenue run-rate increase, 50% enterprise growth, 20 million weekly AI-coding users, and $6.7 billion in Q2 revenue.
Anthropic may move even faster. Reporting says Anthropic could file as soon as the end of August, with ambitions to match or exceed SpaceX’s record IPO scale; Bloomberg and Kimmonismus amplified the timing. That has also revived a mission question: Anthropic’s 2023 Core Views on AI Safety says transformative AI could arrive within a decade while alignment remains unsolved, and Nathan Calvin questioned whether a giant IPO fits that posture. The frontier race is no longer only about models and enterprise contracts. It is becoming a capital-markets race too, where revenue growth, compute spending and investor appetite will start getting judged in public.
🏆 TOP 5 NEWS (Around the Horn)
- Poolside struck a non-exclusive $6 billion Nvidia licensing deal plus a $1 billion investment at a $12 billion valuation before Nvidia’s new cash; Eric Newcomer and Alex Konrad highlighted the unusual structure, including Nvidia offers to 109 employees while Poolside’s founders stay and insist it is neither an acquisition nor an acquihire (buying a company mainly for the talent).
- Stripe agreed to acquire OpenRouter, keeping its brand and roadmap while absorbing the gateway that routes developers across 500+ models from 80+ providers. OpenRouter’s announcement, its model gateway and side-by-side chat playground show the product Stripe is buying; industry context put OpenRouter above 4.5 quadrillion tokens at run rate with roughly 33% month-over-month growth. Stripe’s investor letter, covered by Axios, also declared that “the singularity” began January 1 and cited 41% first-half revenue growth.
- America’s data-center boom is turning into a political problem. Loudoun County, home to 250+ facilities and $1.1 billion in 2025 data-center tax revenue, now requires public approvals for new projects; Tom’s Hardware framed it as one of America’s richest counties hitting the brakes. Separate polling showed roughly 75% opposition, while Jeremiah Johnson highlighted a survey where a nearby coal plant was roughly twice as popular as a data center. Senate Republicans warned that the issue is becoming politically radioactive, The Hill tied it directly to Ohio’s Senate race, Axios traced the power and construction pressures, TechCrunch argued adoption has not translated into public acceptance, and Chamath Palihapitiya warned a broad slowdown could shave 2–3 percentage points from annual GDP growth. There is a reason towns still chase the projects: Crémieux estimated that a roughly 160–200 MW campus could generate about $40 million a year in property taxes, enough to replace the average local property-tax take of a 20,000-person U.S. town. That is his back-of-the-envelope calculation, built from Lincoln Institute property-tax data, its per-capita local-revenue data, an older 50-state commercial-tax comparison, and Turner & Townsend’s data-center cost benchmarks and methodology.
- A randomized experiment with 1,559 Pakistani judges found that a custom generative-AI tool plus targeted training increased district-level case resolution by 6.3% without a clear decline in writing quality or increase in appeals; Kiran Garimella highlighted that the tool and training together mattered more than generic training or tool access alone.
- Generalist GEN-1.5 learned new closed-loop robot tasks from only 3–12 seconds of demonstration with no task-specific retraining, reporting 59% average one-shot success across 10 tasks and 83% after a handful of brief training updates. Generalist and Sholto Douglas circulated the result; Pete Florence connected it to MIT’s 1970 Copy Demo and a 2017 one-shot imitation paper, WIRED showed the robots improvising tool use, and Jim Fan argued the gains come partly from preserving the natural recovery motions and micro-adjustments in human demonstrations while cautioning that the current tasks are still simple.
Honorable Mentions
- Anthropic is preparing customer-controlled retention for Mythos-class models. Reuters reported customers will be able to keep the 30-day safety-monitoring window on their own cloud infrastructure; Boris Cherny said enterprises will own the data while Anthropic retains none, Sholto Douglas said the design still supports longer-horizon cyber monitoring, and Rohan Paul stressed that customer logs remain under customer access controls. OpenAI meanwhile reaffirmed Zero Data Retention for eligible API customers and previewed Private Safety Processing; Axios contrasted the labs’ approaches to safety monitoring and enterprise privacy.
- U.S. agencies warned about attacks on internet-exposed Siemens S7 industrial controllers. The joint federal advisory says attackers are using AI-generated scripts disguised as legitimate industrial tools against manufacturing, energy, water and chemical systems; NSA Cyber amplified the warning, Jason D. Clinton warned of a possible “vulnerability tsunami,” and Cybersecurity Dive covered the campaign.
- California drew a record $366 billion in venture funding this year, more than the other 49 states combined and nearly double its previous high, as AI further concentrated startup capital in Silicon Valley.
- An independent builder pretrained Mini Kimi-K3 on 5 billion tokens scrubbed of benchmark questions for $252.35, with 145 million parameters doing work for each generated token, and reported 33.4% on HellaSwag (a common-sense sentence-completion test) versus 28% for GPT-2 124M. The five-hour Vizuara worklog documents the corpus, decontamination, expert-collapse bugs, training-speed experiments and 24 benchmark checkpoints, while the MIT-licensed GitHub repo publishes the training scaffold. The comparison is specifically against GPT-2 124M on HellaSwag, not a claim that the mini model beats modern 1B-class models overall; the builder says $5,000 is already approved for a larger next run.
🍪 TOP TREATS TO TRY
- ChatGPT Sites builds and hosts websites, apps and games from prompts or local projects, with access controls, analytics, storage, secrets and custom domains; OpenAI Developers showed teammate editing while Codex handles Git and CI, and a follow-up added collaborator editing and customizable Site URLs.
- Claude Academy is a free learning hub for Claude.ai, Claude Code, the Claude Platform, AI Fluency and model limitations; Anthropic’s launch and teaching philosophy emphasize practical fluency and durable human agency.
- OpenAI added transparent-background generation to GPT-Image-2 preview. OpenAI Developers showed the feature, and the official cookbook explains how to generate reusable PNG assets for campaigns, products, mockups and presentations by setting a transparent background and validating the alpha channel.
- Spline V2 rebuilt the browser 3D editor on WebGPU with faster rendering, new lighting/materials, HTML/JavaScript scripting and AI Agent Mode; Spline also added Model Context Protocol support (the standard that lets agents connect to outside tools and data) so outside agents can work with scenes, and the app is live.
- Black Forest Labs’ FLUX Video Upscale regenerates videos at 1080p, 2K or 4K instead of merely stretching pixels. The docs, interactive upscaler and launch blog describe Precise and Creative modes, with supplied pricing of $0.07 or $0.10 per megapixel-second; fal and Replicate also offer it.
- OpenCode made its stealth “Ox Alpha” model free for one week, with a 1M-token context window (roughly how much text and other input it can keep in working memory at once), multimodal input, zero data retention and claimed capacity around 100 trillion tokens per day.
- Notion launched Skills, reusable pages that encode a team’s processes, examples and preferred formats so Notion Agent can load the right playbook automatically or export it as a
SKILL.mdpackage for Claude Code, Codex, Cursor, Gemini or Grok. keb showed one in practice: meeting notes become lightweight whiteboard-style HTML architecture diagrams for before/after discussions, then can be refined in FigJam. - Perplexity’s Agent API gives developers one endpoint for 41 frontier and workhorse models across nine providers, plus built-in web and finance search, fetching and sandboxed code execution; Perplexity Developers framed it as infrastructure for multi-model agent workflows, while Aravind Srinivas argued a real AI developer platform should combine access to multiple model tiers with the tools needed to deploy them into production workloads.
- Grok Build turns one prompt into a published app, game, website, or dashboard with its own domain, while its coding agent can use subagents, a browser, databases and secrets and export projects to GitHub; Grok says Build now ships across web, iOS and Android on every SuperGrok and X Premium plan, and Nick Dobos called its prompt-to-“Add to Home Screen” flow the smoothest he has seen.
🆕 NEW From The Neuron
- Our new interview asks a pretty uncomfortable question: what if we’re spending billions scaling the wrong kind of AI for predicting what happens next? Neuralk CEO Alexandre Pasquiou explains why ChatGPT can summarize business data yet still lose the signals needed to predict what happens next, how tabular foundation models attack that problem directly, and why he thinks they’ll power every enterprise prediction workflow by 2030. He also shows how Neuralk’s Seldon can plug that prediction layer into Claude, ChatGPT, Excel, and AI agents.
🏢 Big Tech & Major Companies
- ChatGPT added an Apple Messages plugin on Mac for ChatGPT Work and Codex, letting users search conversations, catch up, draft or send replies, and combine message context with tasks like checking calendars or flagging spam; a companion post showed the workflow in action.
- AT&T said it is aggressively shifting AI workloads toward open models. One internal-routing account said 40% of employee AI usage already goes to open models, with a 60–70% target and 56% lower coding costs for about a 2% quality tradeoff while handling roughly 45B tokens per day. The Information and Amir Efrati reported a slightly different slice of the program: open models powering roughly a quarter of usage today, a 70–80% ambition, and smart routing cutting costs 80–90% on some applications while AT&T tries to keep Anthropic/OpenAI spend flat.
- Anthropic’s enterprise joint venture Ode with Anthropic, backed by Blackstone, Hellman & Friedman, Goldman Sachs and others, made its first acquisition by buying consultancy Fractional AI to build a forward-deployed engineering team for enterprise Claude implementations.
- Anthropic also made four agent APIs generally available on the Claude Platform. Computer use can batch multiple actions per turn for 20–40% fewer round trips, browser use targets page elements instead of brittle pixels, Skills can be versioned and pinned, and Files support longer retention, 5× rate limits and up to 1 TB per organization.
- Meta has quietly become one of Microsoft’s largest AI customers, reportedly spending hundreds of millions of dollars a year on Azure-hosted models.
- Alibaba saw profit plunge more than 75% to 10.5B yuan after quarterly capital spending climbed to nearly $10B for AI infrastructure, even as revenue rose 9% on cloud growth.
- Google’s Gemma family passed one billion downloads and 100,000 community variants spanning healthcare, biology, space and consumer deployments, and Google launched an Awesome Gemma GitHub collection to showcase the ecosystem. Google for Developers highlighted examples including NASA/Starcloud orbital image analysis, India’s Aarogya Setu health app with 100M+ users, MedGemma clinical tools, C2S-Scale cancer research and DolphinGemma for decoding dolphin vocalizations.
- Nvidia is planning a China-focused inference chip (a chip for running models after training) that uses licensed Groq LPU technology alongside GPUs. The Information, with Reuters also covering the plan, reported small-batch shipments could begin by year-end, several Chinese customers have already placed orders, and the product is being designed to comply with U.S. export controls.
- Waymo built a custom ASIC (a chip designed for one specific job) for its newest robotaxis to improve reflexes and navigation in complex urban environments while reducing dependence on outside chip suppliers.
- Spirit flight attendants challenged Google’s $10M bid for the bankrupt airline’s digital records, arguing confidential employee data could be swept into the AI-training sale; a bankruptcy judge delayed approval after the objection.
- OpenAI reportedly told employees it expects Astra in “a couple of weeks,” with an updated internal checkpoint focused on alignment and reward-hacking; Kimmonismus separately amplified the reported competitive timing around Anthropic’s next model.
- ChatGPT added read-only conversation sharing from ChatGPT Work and Codex on desktop so teams can show process or hand off context without letting recipients edit the original chat.
- OpenAI’s first NVIDIA Vera Rubin racks arrived and are already running the company’s training stack, a milestone in the compute buildout for its next generation of frontier pretraining.
- OpenAI launched AI Futures, a Strategic Futures blog about how transformative AI could reshape power, governance, the economy and individual freedom. Dean Ball framed concentration of power as one of the hardest long-run policy problems and said the blog will feature multiple contributors, not just him.
- OpenAI’s Computer History feature lets the Mac app remember activity across apps and websites while giving users a timeline, privacy controls, app include/exclude settings, pause/clear controls and a way to turn frequently repeated work into skills. The original rollout covered Pro, Business and Enterprise users, while the newer update expanded availability in the EEA, UK and Switzerland.
- Meta’s Muse Spark 1.2 received a second wave of demos and evals. AI at Meta showed visual-to-code generation, chart and knowledge reasoning with tools, spatial reasoning for robotics, bimanual task planning, agentic media generation and detailed captioning used to improve Muse Image/Video training. Design Arena separately ranked it first on Video-to-Website, second on Image-to-HTML and third on Image-to-Frontend.
- Nvidia is also matching customers with Nordic operators that have available land and cheaper power, expanding its role from chip vendor to infrastructure matchmaker.
- Marvell rose nearly 10% after a deal that allows Google to buy up to $12.2B in Marvell shares alongside a multi-year custom-AI-chip partnership tied to Google’s TPU ecosystem (its custom AI accelerator chips); CNBC’s Broadcom analysis examined the competitive pressure on Google’s existing custom-chip supplier.
- Amazon made Alexa+ available on compatible Fire TV devices at no additional cost for U.S. customers, with conversational search, smart-home controls and recommendations; TechCrunch emphasized that Prime is not required.
- Google is giving eligible college students 12 months of its Google AI plan free; Tom’s Guide tied the offer to the company’s broader back-to-school push.
- Meta’s new Mac app is aimed at creators and small businesses, with screen sharing, cross-app dictation and tools for analyzing Meta ad performance and producing reports/content; The Verge highlighted the general desktop-assistant experience.
- GitHub’s August 17 outage lasted 7 hours 47 minutes after a critical Central U.S. infrastructure component failed to scale under peak traffic, disrupting authentication, Actions, APIs, pull requests, issues and Copilot. GitHub says it is accelerating capacity toward 3M+ CPU cores and 120 PB of storage, continuing an Azure migration that already carries 58% of load, isolating critical systems and standardizing retry budgets to prevent retry storms.
- Apple Music will begin labeling tracks “materially generated using AI” later this year, requiring providers to disclose AI use in artwork, audio, composition and music videos so those labels can be surfaced to listeners.
- Supermicro’s board said it found no evidence CEO Charles Liang knew about an alleged $2.5B Nvidia-chip smuggling scheme led by co-founder Wally Liaw, though DOJ, SEC and Taiwanese scrutiny continues and Liaw’s trial is scheduled for March 2027.
- Ramp’s data suggests the OpenAI/Anthropic enterprise race flipped again in Q3. Ara Kharazian, using Ramp AI Index transaction data from 70,000+ U.S. businesses, said OpenAI’s quarter-over-quarter enterprise growth reached 82 versus Anthropic’s 76, attributing the reversal to GPT-5.6 Sol developer adoption versus Fable 5 pricing and retention concerns. dax argued the contrast reflects a deeper strategic split: OpenAI tolerates more pain to win bottom-up users while Anthropic prioritizes a safer enterprise posture, and bottom-up growth has historically won despite carrying more risk.
- Amazon is expanding Prime Air drone delivery to nearly 500 U.S. cities and towns by the end of 2026, with packages under five pounds arriving in as little as 30 minutes; CNBC reported a top executive expects one million deliveries this year.
💼 AI Productivity, Labor & Economics
- SK Hynix reached a tentative labor agreement including a 6.3% salary increase and a shift in performance bonuses toward 40% cash and 60% shares; Tom’s Hardware reported a potential $1.79B profit-sharing pool worth roughly $50K per employee.
- Goldman Sachs found AI is already weighing on employment in developed economies, with the clearest effects in call centers, software publishing, consulting and advertising, especially among entry-level workers.
- A LinkedIn study found women represented only 26% of U.S. hires into high-paying AI jobs and 13% of leadership roles at AI firms globally while being over-represented in clerical work most exposed to automation.
- The AI boom may also make ordinary products pricier: The Atlantic argued the AI-driven memory shortage is starting to hit cars as modern vehicles depend on centralized, high-memory computers for displays, driver assistance and software updates.
- Consumers surveyed about financial advice said they want AI to support financial decisions rather than replace human advisers.
- Phin Barnes argued founders should resist VC pressure toward capital-intensive deep-tech moonshots that maximize dependence on capital markets and instead favor capital-efficient software that distributes AI intelligence into science, medicine and energy.
- Gabriel argued large enterprises ultimately ask AI to do three things: cut cost, preserve or improve quality, and prove both objectively. Replies added that basic adoption is still a bigger concern for many CEOs than “accelerate R&D,” while mid-market firms can often move faster with less red tape.
- Aakash Sabharwal argued the hard enterprise-AI mapping is Business KPI → Use Case/Task → Agent Eval, with information loss at every layer. Bespoke workflows go stale quickly, subject-matter experts have little time to label, and automated evals miss domain nuance, giving data companies with contributor networks a role in keeping evaluations fresh.
- Ryan Carson described an agent-era hiring process at HelloUntangle where candidates submit an unedited full-screen recording of shipping a real feature with an agent, finalists get paid Devin access to deliver a merge-ready PR on the actual repo within 16 hours, and hiring decisions rest on the work and agent threads rather than meetings.
- Stanford’s GDP-B Surplus Observer estimates the consumer benefit from goods and services that normal GDP misses. The Digital Economy Lab says generative AI was creating roughly $183B in annual U.S. consumer surplus by July 2026, up 58% year over year, based on willingness-to-accept surveys.
🤖 AI Agents & Infrastructure
- SPADE trains one model as both an environment designer and a reasoning agent, creating executable challenges near the agent’s current ability frontier; the team reported a +5.3 average gain over the strongest fixed-environment baseline across eight held-out benchmarks at 30B scale. Bo Liu introduced the work, a follow-up added context, and the team released live demos, paper discussion, code and models.
- Dots Studio open-sourced dots3-note Preview, a 280B-total/16B-active Mixture-of-Experts model (only a slice activates for each request) with a 512K context window and text, vision and speech support. Its long-horizon agent design uses TEMPO for self-critique and credit assignment, online memory updating and recursive improvement; Omar Saravia highlighted the release, and a free OpenRouter endpoint is available.
- Claire Vo listed practical browser-agent jobs ranging from closing the books and clearing Gmail or LinkedIn inboxes to configuring SaaS tools without APIs, running browser QA on code changes, completing security questionnaires, registering for conferences and cancelling subscriptions.
- James Broughel argues “untethered” AI agents with no accountable human or institution may need chokepoint regulation through intermediaries such as compute providers and payment processors; his thread compares unregistered agents to stateless vessels refused by every port while distinguishing that approach from past political abuses of Operation Chokepoint.
- Binance Agent OS connects agents built with ChatGPT, Claude Code, Cursor and other tools to market data, trading, payments and on-chain activity. Binance says users retain controls through subaccounts, blocked withdrawals by default and approval gates, while its launch announcement positions Agent OS as a standardized financial-access layer for AI applications.
- Ramp Router gives U.S. users one API for models from OpenAI, Anthropic, DeepSeek, Moonshot, MiniMax, Nvidia, xAI and others, with routing rules for cost, benchmark performance and hard-problem escalation plus spend/latency dashboards.
- A local Hermes-agent demo on NVIDIA RTX Spark showed a persistent agent monitoring communication channels, ranking a broken booking site as urgent, loading the entire codebase into a local Qwen model on 128GB unified memory, finding the root cause, rebuilding the site, launching a computer-use agent to QA the UI, merging the fix and continuing work from phone messages while the developer stepped away.
- Claude Developers released a cookbook for pairing Claude Managed Agents with AG-UI (an open protocol for streaming agent activity into interfaces). The Anthropic quickstart maps each chat thread to a managed session and demonstrates a finance assistant that streams text/tool activity and renders payoff timelines, growth projections and budgets inline.
- The Harness Continual Learning paper, highlighted in Omar Saravia’s summary, treats the agent harness itself as the thing that learns around a frozen foundation model: prompts, memories, tools, skills and routing evolve over time. It identifies “harness-level forgetting” and uses a Continual Optimizer plus Evaluator to commit updates only after checking improvement, retention and validity, reporting more than 10% relative gains across textual, multimodal and open-world tasks.
- Omar Saravia separately argues the highest-value agent workflows are collaborative: humans verify hard outputs and encode that judgment into reusable skills or verifiers. His thesis is that domain expertise and taste become more valuable, not less, because verified human judgment compounds into the system.
- Every’s Monologue team has turned specialized Codex agents into a virtual engineering team. Every’s write-up describes one human engineer routing work to agents with isolated AGENTS.md files, skills, memory and codebases; the accompanying post highlights ticket routing, bug fixes, pull-request creation (proposed code changes) and GPT-5.6 context handoffs across projects, shifting the human job toward prioritization and review.
- Lindy is positioning itself as an AI teammate connected to 1,000+ tools, with editable plain-file memory and the ability to research, produce reports, join meetings, run recurring briefs and custom skills across Slack, iMessage and email. Its Chrome-extension launch adds inbox prioritization, style-matched draft replies and learned labeling.
- Sanja Fidler co-founded Veeda AI to build high-fidelity simulated reality powered by world models so robots can learn through safe, scalable interactive trial-and-error rather than only expensive real-world data.
- Grok Bot, from the Cursor + SpaceXAI team, is already being used for mundane parental mental-load jobs with almost no setup: scanning school emails, finding forgotten gift cards, booking doctor/hair/massage appointments, listing Marketplace items and ordering school photos.
- Foundation from Chroma learns from agent sessions across Claude Code, Codex, Cursor, Slack and other tools, then maintains a versioned, hyperlinked, provenance-tracked wiki of what a team knows. Jeff Huber announced the SOC 2-compliant research preview, which supports real-time sync, continuous improvement and team knowledge that updates as agents work ($30/user/month in the supplied product context).
- Hermes Agent Cloud offers one-click deployment for an always-on autonomous agent that remembers what it learns, runs 24/7 while scaling to zero when idle, connects through Telegram, Discord, Slack, email and CLI, schedules work in natural language and executes in isolated sandboxes.
- DreamGym synthesizes agent experience instead of paying for every real-environment rollout: a reasoning-based experience model, a stored bank of past examples and an adaptive task generator create increasingly useful synthetic interactions. Zhaorun Chen and collaborators report more than 30% gains over baselines on WebArena and results competitive with GRPO/PPO, two common reinforcement-learning training methods, while relying on synthetic interactions for training.
- Vendo is an open-source embedded-agent layer that lets B2B SaaS users describe the dashboard, workflow or mini-app they need; the agent acts through the host product’s API as the signed-in user and generates sandboxed React views that can be pinned, shared or triggered. The Launch HN thread frames it as a way for customers to build the long-tail features that otherwise sit on a roadmap.
- Matt Pocock’s /wayfinder skill gives agents a decision map, research/prototyping tickets and evolving sessions for projects where the plan is still fuzzy; swyx and Latent Space highlighted the workflow.
- Qinyuan Ye and collaborators found memory-based self-improving agents can be surprisingly fragile: noisy evaluations get amplified during self-improvement, performance changes substantially depending on task order, and underspecified tasks or environments can make apparent progress unstable. Rubrics and feedback helped but did not eliminate the problem, so the authors recommend multiple-run reporting, order stress tests, and interfaces that preserve human oversight.
- Chaokun Chang and collaborators found agent workloads behave very differently from ordinary chatbot inference: non-model components dominated latency in half of ten tested apps, individual sandboxes peaked around 28 GB of memory, and heterogeneous resources produced up to 32× latency differences. Their task-aware scheduling cut latency 29–40%, communication-aware placement improved performance up to 4.5×, state offloading reduced memory 4.6×, and tool-result caching removed 35.2% of redundant searches; DAIR.AI highlighted the same result as evidence that serving agents increasingly means optimizing everything around the model, not merely model inference.
💻 AI Coding & Developer Tools
- Salesforce introduced Slack Code, dedicated code channels where humans and agents from Anthropic, GitHub, Cognition, OpenAI and Vercel can write, review diffs, inspect previews and ship together. Marc Benioff pitched it as live multiplayer agentic coding, while Nick Dobos called the direction “MMO vibecoding.”
- Claude Code added a Concise output style that leads with the result and stays short by default; creator Boris Cherny called it a quick band-aid while Anthropic works on a longer-term fix.
- A developer’s Mac file-system deep dive benchmarked agent-heavy worktrees, pnpm installs and small-file deletion across APFS, ext4, XFS, ZFS and Btrfs. The author found APFS dramatically slower on those workloads, then reported much faster worktrees/installs and roughly 44% storage savings after moving agent development to Linux with XFS plus VDO/LZ4-style compression; his broader point is that modern parallel agents turn filesystem behavior into a visible developer-productivity bottleneck.
- Cursor’s cloud-agent update, announced on Cursor’s account with Matty P’s note, lets always-on agents subscribe to pull requests, Slack and schedules, drive work to completion, spawn isolated-VM subagents, pin skills as permanent Custom Modes, use
/goalfor long-lived objectives and accept steering without interrupting the current action. - Chris McCormick started writing model code as one explicit
forward_backwardfunction with no autograd ortorch.nn, gradients on.gacc, and manually managed activations, sometimes recomputing cheap operations like ReLU instead of storing them; Claude wrote the backward math. Andrej Karpathy argues agents make this teardown of abstractions newly attractive because they can handle the math, verification and drudgery, and his follow-up extrapolates toward a simple scalar-Python/microGPT-style specification as the durable artifact, with frameworks becoming compiler-like intermediate representations. - Ethan Mollick notes that ChatGPT Work/Codex/Chat and Claude Cowork/Code/Chat now have different plugins, skills, memories, permissions and files, making it hard to know which mode can do what. Nick Dobos adds that apps fail to surface a clear inventory of available capabilities, so switching modes feels like amnesia.
- Allie K. Miller published a cost-tiered recommendation for business users choosing between Codex, Claude Code/Cowork and ChatGPT, favoring Codex at many budgets for feature completeness, live voice, image generation, token efficiency, easier tool connections and scheduled tasks.
- Nathan Habib announced Long-Horizon Terminal-Bench, a contamination-resistant 46-task test of whether agents can survive 300+ terminal steps without losing state, using hidden verifiers to score long-running work.
- Modular open-sourced Mojo, including the full compiler, tooling and related components, under Apache 2.0 plus LLVM exceptions after the 1.0 release; the standard library was already open. Phil Eaton highlighted that developers can now build and distribute binaries without depending on closed compiler pieces. A Hacker News discussion pushed back on the broader Python-replacement narrative and focused on Mojo’s MLIR and machine-learning-kernel roots.
- Unreal MCP is now in beta inside Unreal Editor for Fortnite, letting Claude Code, Codex and Cursor write and iterate Verse, place/configure devices, edit Scene Graph entities and start or debug play sessions directly in the editor; Fortnite Developers demonstrated the workflows.
- Anthropic’s early
/designcommand for Claude Code reads a codebase, matches the existing UI style, produces shareable artboard options users can edit, and can then implement the selected design back into the project. - Devin can now work inside dedicated Slack Code channels. Jeff Wang showed Devin creating collaborative spaces for complex work, surfacing diffs and pull requests (proposed code changes), testing across platforms and proving results with video while staying quiet by default; Cognition announced the integration, and its Slack Etiquette write-up details a year of prompt, harness and Slack API changes built around “silence is the default.”
- Hamel Husain and Shreya Shankar’s AI Evals for Engineers & PMs course teaches teams to trace agents, turn vague failures into reproducible cases, build trusted LLM-as-judge and code evaluators, wire regression tests into CI/CD (the automated testing-and-deployment pipeline), red-team for safety and optimize cost while improving accuracy across a real agent project over 17 live sessions ($4,200).
- oMLX reached 20,000 GitHub stars six months after its 0.1.0 release, with Jun Kim crediting 200+ contributors for building a local MLX app that tries to preserve both convenience and user control.
- Nick Baumann shared that people at OpenAI are using Codex as a continuous Spotify/Apple Music DJ that maintains a vibe-based queue, and he built a small Mac app to feed it reactions and playlist suggestions.
- PyTorch highlighted IBM Spyre work where AI coding agents wrote runtime adapters (small compatibility layers) for unsupported operations and memory-alignment issues, enabling 7,960 of the top 10,000 Hugging Face embedding models on a new accelerator; 6,804 passed full end-to-end tests. The PyTorch write-up argues this can collapse months of specialist model-enablement work into a much faster day-one workflow.
- Huzzah is an experimental AI coding editor where the durable source is concise, declarative
.hzpseudocode. Saving generates the real implementation, while later edits send only the diff back to the model so the human-readable intent stays versioned instead of disappearing into chat history. - Vomit runs Claude 5’s verbose output through a separate local OpenAI-compatible model that rewrites it into clearer English while trying to preserve intent; the Hacker News discussion centers on frustration with model prose that sounds authoritative while making simple work harder to parse.
- Tidal Cycles is a free Haskell-based live-coding environment for composing algorithmic sound, notes and parameters in real time with SuperCollider; a Hacker News thread resurfaced the project for programmers interested in making music as code.
- Niki found that scroll-event burstiness and memory patterns alone could train a lightweight machine-learning classifier (LightGBM) to distinguish humans from scraper traffic at about 73% accuracy on single-page visits, a lightweight bot signal for pages that require scrolling; the Hacker News discussion dug into the measurement tradeoffs.
- Cursor’s Continuity treats a write-ahead log (an append-only record of changes) in S3 as the source of truth for Git storage, materializes repositories as warm caches and lets any node act as primary. Cursor reports linear read scaling to 100 replicas and high push throughput without the consensus or packfile bottlenecks of earlier designs; Hacker News compared the architecture to familiar database WAL-and-compaction patterns.
🔬 AI Research & Models
- Superwhisper released S1-mini, its first open-weights 0.6B-parameter model for on-device transcript processing. A companion launch post described it as a local text normalizer that removes filler words and false starts, fixes punctuation and capitalization, and converts spoken numbers, dates, currency and email addresses; the model is available on Hugging Face.
- Patronus AI open-sourced FigmaTrace, 3,469 trajectories covering 200+ hours of real human Figma work across 10 creative skills and 126 long tasks. Patronus reports up to a 26% improvement in out-of-domain step-level actions and says smaller open models can match or beat frontier models on agentic GUI tasks; the announcement, FigmaTrace paper and dataset are public.
- Developer Hyeonseok Jung built Visionary, a 300M-parameter world model for the SO-101 robot arm that learns rigid-body, shelf/door and deformable-object physics and can produce stable long rollouts in-browser.
- Rei Labs exposed Temporal Context Projection and Counterfactual Utility Plasticity in its Adapt-1 Preview API, mechanisms meant to connect present decisions to earlier events and learn which internal predictors help or interfere. Rei released replication experiments and announced the update.
- Anil Seth highlighted a Physics of Life Reviews paper arguing recurrent feedback loops across cellular, local and global brain scales may be the common implementation primitive underlying major theories of consciousness.
- GitLab co-founder Sid Sijbrandij described going “founder mode” on osteosarcoma after standard care was exhausted, assembling a personal advisory board, generating more than 30TB of multi-omics data and pursuing parallel experimental therapies including a personalized mRNA vaccine and radioligand treatment; his post says he has had no evidence of disease for more than a year.
- Token Gremlin alleges Anthropic’s recent Opus quality problems may stem from attempts to force textual watermarking into outputs, blaming the change for hallucinations and degraded writing. That causal claim remains commentary, not an established finding.
- Doron Zeilberger argues AI has become normal enough that we should drop the “A,” call machine intelligence simply “I,” and rename human intelligence “NI” for natural intelligence. Not sure if this is sarcastic or not...
- Meta’s WildArtifactBench Bird Sound task asks agents to identify 46 bird species from audio plus YAML descriptions and produce a perfect mapping file without pre-existing bird libraries. AI at Meta described it as one of 10 released tasks from a broader internal framework for evaluating multimodal agents on messy real-world work via human/agentic preference judges.
- Harvey introduced Tenet, its first post-trained legal model built from Kimi K3 with Fireworks, reporting an 82% lift in all-pass rate on LAB and 22% on LAB Contracts versus the base model, SOTA performance on LAB Contracts, second place on LAB and less than one-quarter the cost of frontier models. Gabe Pereyra detailed synthetic/public/human-expert training plus specialist subagents; Mercor reported a +15.2-point lift on its external Corporate Law evaluation to 74.0%, while the benchmark contains 160 long-horizon tasks authored by practicing Big Law attorneys.
- diffusion-model foundations paper unifies continuous diffusion for images with discrete/categorical diffusion for sequences, derives a shared training objective and explains how forward noise determines reverse dynamics; Vincent Pauline announced updates after community feedback.
- KernelBench is an open benchmark for agents writing GPU kernels. Elliot Arledge highlighted DeepSeek V4 Pro reaching 9.7% of theoretical peak on a GLM-5.2 fused Mixture-of-Experts task versus Opus 5’s 10.7%, while also noting DeepSeek correctly exploited an unstated prompt constraint by using cuBLAS.
- Liquid AI released LFM2.5-DSpark, lightweight speculative-decoding draft models that propose token blocks for the larger target model to verify together. Liquid’s launch post reports up to 3.18× throughput on Nvidia H100 data-center GPUs, 2.87× on M4 Max and roughly 50% lower latency on tool-calling workloads with identical greedy outputs. Draft checkpoints are available for 1.2B-Instruct, 2.6B and 8B-A1B; SGLang reports 2.1–2.4× H100 speedups and llama.cpp about 2.1× on its integration. OpenMed called tool-heavy clinical workflows a strong use case. Developer Abdur Rahim’s mlx-dspark extends DSpark/DFlash to Apple Silicon, with his post reporting up to 4× lossless decoding across several model families; Liquid’s Ramin highlighted similar on-device gains for LFM models.
- Google uploaded TIPS v1-g14, a vision-and-text encoder combining contrastive learning with masked-image modeling to learn spatially rich features for dense tasks such as segmentation and depth estimation. Hugging Papers resurfaced the model and its TIPS paper, originally associated with the 2024/ICLR 2025 work.
- Datapoint AI released more than 2.16M validated pairwise human image-preference votes covering all 435 pairings among 30 state-of-the-art text-to-image models, 500 prompts across 10 categories, controlled fixed-seed generation, no prompt rewriting/best-of-N and annotator trust scores. The dataset, Image Bench and $1M research-grant program are public; the dataset also holds out 50 prompts for clean evaluation.
- The AI post-training paper, summarized by Omar Saravia, found agents tend to lock in a high-level strategy at the first step and spend the remaining budget making local adjustments instead of spontaneously reconsidering it. Experience scaffolds improved execution (+12.6 on GSM8K (grade-school math) and +40.8 on HumanEval (Python coding problems) in the supplied summary) without fixing strategy lock-in; human guidance could redirect the opening choice before agents fell back into local loops, and extra compute helped easy tasks much more than the hardest ones.
- Abra, highlighted by Sway, studies scaling laws for text-to-image diffusion from 60M to 2B parameters and finds compute-optimal training around 200 image tokens per parameter, roughly 10× the Chinchilla-style ratio cited for LLMs, with predictable scaling across generative quality and representations.
- LifeGPT learns Game-of-Life state transitions without being told grid size or boundary conditions and can recurse its predictions; AutomataGPT, with a published paper page and commentary from Jaime Berkovich, extends the idea to learning and recovering rules across many two-dimensional cellular automata.
- Ornith released the open-source Ornith-1.5 family: 9B dense, 35B-A3B Mixture-of-Experts and 397B MoE models trained through an end-to-end self-improvement loop in which the model proposes tasks, builds scaffolds and generates reinforcement-learning rollouts. Ornith reports state-of-the-art open-model results at comparable sizes and Claude-Opus-level performance on some agentic/coding tests; the 35B-A3B checkpoint activates about 3B parameters per token, is MIT-licensed and is available in quantized variants.
- AI in spatial pathology is using deep-learning systems to segment/classify cells in whole-slide tissue images, extract spatial neighborhood features and power foundation models for biomarkers and pan-cancer detection with less task-specific labeled data.
- Qwen3.8-27B is an Apache-2.0 27B vision-language model with a gated DeltaNet-plus-attention architecture, 262K native context that can extend to 1M, support for images and hour-scale video, and configurable thinking effort. Qwen reports 61.7 on SWE-bench Pro (real software-engineering tasks), 90.3 on LiveCodeBench (coding problems) and 84.3 on OSWorld (computer-use tasks). Benjamin Marie found the xhigh mode substantially more token-efficient than Qwen3.6 on most non-agentic tasks, suggesting complaints that it “thinks too much” may be concentrated in agentic workloads; his full benchmark write-up compares accuracy, token use and KV-cache memory against Muse Glimmer.
- Agape Keleta says he trained a model that removes stereotypical AI prose by shifting sampling toward higher-entropy human micro-style while preserving overall coherence, and that it also strips Claude’s statistical watermark; he reports Pangram v4 classified every output in his evaluation set as human-written.
- Hannah Chung demonstrated learned KV-cache transfer, where an architecture-compatible Qwen 8B model prefills a prompt and the cache is adapted for Qwen 32B. She reports 17.8% lower time-to-first-token, $641 savings per 1M requests, 82.5% held-out top-1 agreement and 91.9% recovery of prompt evidence versus NVIDIA’s method.
- Personalized mRNA cancer vaccines kept producing unusually strong signals. Yannick Buccella argued the technology crossed from sci-fi toward clinical reality in 2026, citing melanoma Phase 3 results around 49% lower recurrence risk plus early pancreatic data and trials in kidney, lung, bladder and brain cancers. He stressed that the current sweet spot is post-surgery prevention of recurrence, manufacturing still takes weeks and can cost $100K–$300K, and the first approvals are more likely in 2027 than immediately. Dr. Singularity highlighted a separate approach aimed at training immunity before pancreatic cancer develops, and Trung Phan amplified the Moderna/Merck Phase 3 melanoma result after the earlier Phase 2 recurrence reduction.
- MiniMax-H3-RAVEN-Streaming-LoRA is a preview rank-128 LoRA adapter (a small set of add-on weights that changes a model without retraining the whole thing) that turns MiniMax-H3 into a causal streaming video generator, producing autoregressive chunks with a sink of 2 and window of 2 at 768×1376/24fps instead of generating a full bidirectional clip at once. AI Search highlighted the real-time potential, while MiniMax praised the community project but said true real-time generation still needs more inference acceleration and invited builders to an H3 hackathon.
- Pietro Barbiero introduced a Penrose-inspired graphical notation for interpretable neural architectures with global model views, geometric diagrams of every tensor operation and a one-to-one mapping to PyTorch
einsumcode. His Graphical Design paper formalizes the notation across concept bottlenecks, sparse autoencoders, prototype networks, neural additive models and mixtures of linear models. A separate Guide Labs paper argues that making interpretability a training-time constraint rather than a post-hoc tax can improve disentanglement and human-concept alignment with scale; Guide Labs built the Steerling-8B diffusion language model to attribute any output span to input tokens, concepts and training data, steer concepts in a closed loop without retraining, and remain competitive with peers trained on 2–16× more compute. - Single-Rollout Asynchronous Optimization replaces GRPO-style group sampling with one rollout per prompt, adds value-model training and strict two-sided token clipping, and reports stable 1,000-step training plus gains on SWE-Bench Verified, BeyondAIME and IMOAnswerBench; Zixuan Li highlighted that it was deployed for GLM-5.2.
- Jyo Pari studied dynamic compression for in-context continual learning, where a model can revisit and reorganize its earlier state as it discovers what needs to be reused instead of writing each token once into a fixed-size recurrent state. The work is currently in synthetic regression, so whether it scales to natural language remains open.
- Self-Organising Digital Circuits extends Neural Cellular Automata from grid patterns to functional logic: a topology-masked Transformer configures Boolean-gate lookup tables so circuits can self-assemble from noisy states and reroute around permanent or soft faults. The arXiv paper reports near-perfect soft-error recovery and generalization to larger unseen graphs; Sebastian Risi announced the work.
- Verifier Frontier tests how small an AI verifier can get. Yannick Detrois found exact-task verifiers for Countdown and Maze could work at roughly 0.63M parameters after about two minutes of pretraining from scratch on one H100, while 1–2M-parameter models could handle faithfulness judging and match or beat much larger systems, including zero-shot Gemini 2.5 Flash on that task; his announcement emphasizes how tiny task-specific evaluators can be.
- HydroGym is a solver-independent reinforcement-learning platform with 60+ validated 2D/3D fluid-control environments spanning laminar to turbulent flows, Reynolds numbers up to 4×10^5 and varying Mach numbers. Agents repeatedly discovered reusable control ideas such as boundary-layer manipulation, acoustic-feedback disruption and turbulent-wake reorganization; agents trained only in cheap surrogate environments then transferred zero-shot to a 3D wing, cutting local skin friction 38% while reducing exploration costs by four orders of magnitude. The authors note the breadth of that generalization is still open. The code is open, and Steven Brunton highlighted the result.
- Jie Tang argues parameter count only makes sense alongside data volume, compute allocation and inference conditions, tracing the path from Kaplan’s over-parameterization mistake through Chinchilla, over-training and Mixture-of-Experts. He points to GLM-5.3 using the same base as 5.2 plus a month of long-horizon reinforcement learning as evidence that post-training becomes a major scaling dial once a knowledge threshold is reached. Liam Fedus recalled Switch Transformers’ extreme 1-of-2048 routing, with 1.6T total parameters but fewer than 3B active, as excellent at knowledge yet “dumb as bricks” on reasoning, reinforcing his view that compute (the amount of math the model performs) drives intelligence while parameters drive knowledge and that the ideal tokens-per-parameter ratio depends on the task.
- Gerard Sans argues apparent “broken scaling laws” are real rather than curve-fitting artifacts because global loss averages treat noisy samples equally; he wants new mathematical objects that track local model-family envelopes and data composition instead of one entangled average.
- Nathan Godey’s Matryoshka LM Suites trains nested submodels with independently adjustable width and depth in one run instead of training every model size separately. The Matryoshka paper reports the same performance at much lower training compute, free online distillation between nested models, better mutual alignment and 10–30% faster speculative decoding because the suite can share KV cache, memory and activations; a 3B model is public. The work extends the earlier MatFormer, which first showed how nested Transformer components could yield elastic model sizes from a single training run.
- TRACES proposes an evaluation paradigm for “Discoverative AI” that scores how rigorously models investigate open problems with no answer key rather than whether they retrieve known answers. Apodex launched it with six scored capabilities covering Tools, Repair, Alternatives, Coherence, Evidence and Scope, a technical report spanning biomedicine and frontier engineering, and open calls for both new problems and solver systems.
- Google DeepMind’s Recirculation feeds a small amount of activity from deeper Transformer layers (the stacked processing stages inside most LLMs) back into shallower layers at inference time so a frozen model can keep updating an internal belief state without changing its weights. On Gemma 3 models, the team reports 23% lower perplexity and a 21% GSM8K accuracy gain without retraining or added generation latency; Turing Post highlighted how it differs from chain-of-thought or simply looping the model.
- DiffusionGemma turns an existing Gemma 4 Mixture-of-Experts checkpoint (a model where only part of the network activates for each token) into an experimental discrete-diffusion language model that refines 256-token blocks in parallel, reaching roughly 1,500 output tokens per second on one H100 while retaining thinking mode, multimodality and long context. A Hacker News discussion focused on the fact that the team converted an existing checkpoint instead of training from scratch.
- The DESI Legacy Imaging Surveys team released the largest 2D universe map: a 5.6-trillion-pixel public map covering about 75% of the sky and nearly four billion celestial objects, built as a foundation for 3D dark-energy studies and other astronomy. Hacker News resurfaced the public dataset this week.
🏛️ AI Policy, Governance & Security
- Greg Brockman argues the OpenAI-Hugging Face cyber incident showed agents can chain unknown vulnerabilities and leaked credentials into real attacks, creating a temporary “defender’s window” in which organizations should use AI to find and fix vulnerabilities before open-weight cyber-capable models narrow the advantage.
- CTIFoundry reorganizes cyber-threat intelligence into a deterministic graph connecting CVE, CWE, CAPEC and ATT&CK data plus grounded report passages, then exposes that structure to agents through seven tools and three procedural skills. The authors report +0.19 to +0.28 F1 gains on CTIConnect, with a smaller model using the structured scaffold beating a flagship model working from flat documents while using roughly half as many tool calls; DAIR.AI highlighted the result as evidence that organizing knowledge at indexing time can matter as much as upgrading the model.
- China has been restricting or delaying exports of germanium, quartz-based products and neodymium magnets to Taiwan since 2025, creating bottlenecks for optics, photonics, semiconductor equipment and aerospace suppliers; Digitimes reports delays have already cost some Taiwanese suppliers orders.
- Saif M. Khan launched the nonpartisan Center for Technology & Statecraft; its founding agenda focuses on technically grounded work around automation/social-contract policy and long-run U.S.–China AI competition and stability.
- Nick Caputo launched Model Constitution, a publication and proposed public-good “Philadelphia Project” for studying company AI constitutions and eventually drafting a model constitution compatible with democratic checks and balances.
- Elizabeth Sherwood-Randall argues AI plus biotechnology is lowering barriers to synthetic-pathogen design for non-state actors, making classic deterrence less effective and requiring “deterrence by resilience” built on early detection, surge manufacturing and global norms.
- A separate AI-biology security report argues the federal government needs more expertise and stronger participation in deterrence and coordination with outside groups.
- The American Medical Association released a framework insisting physicians remain in the loop for clinical judgment, human connection and technology stewardship rather than treating AI as a full substitute for medical practice.
- Lockheed Martin argues its AI Center, testing processes and partnerships are intended to deliver trustworthy, mission-ready systems “at the speed of relevance” for national security.
- The Trump administration launched a $5 billion Genesis Mission aimed at using AI to accelerate U.S. scientific discovery, with the administration framing it as part of a new “golden age of science.”
- tenobrus urged Anthropic employees to push leadership for a public parallel pause matching OpenAI’s frontier reinforcement-learning halt, arguing that a costly visible signal of willingness to coordinate despite rivalry is exactly the kind of foundation real safety cooperation needs.
🛠️ AI Tools & Products
- Google DeepMind launched Backstory, an experimental Gemini-based tool that investigates an online image’s origin, prior appearances, edits and possible AI generation. Nieman Lab reports fact-checkers are combining provenance checks, reverse-image search, SynthID/C2PA signals and historical context into cited reports in minutes; its post highlighted the workflow.
- Meta Pocket expanded to U.S. users, letting people generate and share interactive phone games from prompts that can react to touch, tilt, sound, camera-roll photos and the live camera.
- Soniox Text-to-Speech, announced on Soniox’s account, generates expressive speech in 60+ languages with instant voice cloning, low-latency streaming, character-level timestamps and controls for difficult names/alphanumerics/medical or finance terms.
- Google added five study tools across Search, Lens and AI Mode: interactive concept visuals/simulations, custom practice quizzes with explanations, Lens step-by-step problem solving from a photo, notebooks for source-grounded study, and auto-generated one-pagers/slides/spreadsheets from notes.
- Google expanded generative UI from AI Mode into AI Overviews, letting Search dynamically assemble custom visual layouts, interactive tools and simulations around a query instead of returning only a static block of generated text.
- Google Flow gave three creatives unlimited access to its AI creative studio to build campaigns for local organizations, producing a regenerative-farming story, a 200-year ferry-history film and a caregiver-support film built around a mirror metaphor.
- Wispr Flow was the subject of a deliberately skeptical NYT experiment in which Amy X. Wang dictated an entire column through the app and found the unedited output imperfect enough to want to rewrite it heavily despite the tool’s productivity pitch.
- Something Big is Matt Shumer’s new weekly AI newsletter promising clear high-signal takes and practical workflows; its launch issue recounts a controlled real-world experiment in which an LLM called Luna acted as a middle manager and fired an employee.
- Artificial Analysis provides independent comparisons across model quality, cost per task, output speed and latency for many frontier and open models.
- Artificial Analysis’ image tools now make model choice more concrete. Its Qwen-Image-3.0-Pro update put Alibaba’s model at #6 for image editing after an 83-point Elo jump (the chess-style head-to-head ranking score) and #9 for text-to-image after a 48-point Elo jump, highlighting realism, precise small text and support for prompts up to 4.5K tokens across 12 languages; the faster Qwen-Image-3.0 also climbed sharply. The Text-to-Image Leaderboard ranks OpenAI GPT Image 2 (high), Reve 2.1 and Google Nano Banana 2 among the current leaders, while the Image Editing Leaderboard is led by Microsoft MAI-Image-2.5-Pro, Reve 2.1 and OpenAI GPT Image 2 (high). Image Arena is the blind head-to-head voting interface that feeds those Elo rankings and also exposes generation speed and price.
- MoCHi is a personal ChatGPT-style workspace with private chats, bring-your-own-provider keys, memory, personas and multi-model support. The creator’s V2 release focused on speed and memory fixes, a follow-up explained the new memory behavior, and a later cost tip suggested shrinking the context budget when API bills climb.
- George Kenwright demoed Gemini 3.7 Flash inside Antigravity generating an Omni video that plays natively inside Google Sheets, Calendar and Chat, including frame edits, a useful glimpse of generative media moving into ordinary productivity surfaces.
- RollTab’s creator trained a 125M-parameter Transformer that autocompletes live piano performances from a short MIDI prompt entirely on-device, running at about 108 notes per second on an iPhone 15; the Show HN launch describes it as Copilot for a piano and says the app is free.
- Zoneless is an Apache-2.0 Stripe Connect alternative for global marketplace payouts over stablecoins (crypto tokens designed to hold a stable value), with no platform fee and an API/webhook shape designed to run beside Stripe or PayPal. The creator’s Show HN post reports 5,000+ sellers onboarded, 3,000+ payouts and 74% of new sellers on his marketplace choosing it over Stripe.
- HyNote captures meeting system audio without sending a bot into Zoom, Meet or Teams, transcribes on-device on Mac, and turns meetings, audio, PDFs, YouTube links, images and web clips into searchable summaries, action items and a private knowledge base; its Product Hunt listing says it is free to start.
- MiniMax Design H3 is a multimodal creative agent that decomposes a brief across image, video, voice and editing models, assembles commercial content on a node-graph canvas, keeps local files on the user’s machine and saves proven workflows as reusable Skills; MiniMax’s Product Hunt page provides the broader product context.
🤖 Robotics & Autonomous Systems
- China deployed nearly 50 SUPCON robocops across roughly eight cities for daytime traffic duty. The 1.88-meter robots use cameras and radar to identify violations, direct traffic and answer questions, but have no arrest powers or weapons and remain centrally monitored.
- Uber launched autonomous rides in Zagreb with Verne and Pony.ai, making the Croatian capital the first European city where users can book a self-driving vehicle in the Uber app; initial rides include a safety operator.
- China’s humanoid-robot market is getting a meaningful share of revenue from government-backed training centers that collect teleoperation data and sell it back to robot makers, creating a circular policy-supported demand loop that complicates how much end-user demand is truly commercial.
- Unitree’s $1,600 Go2 traces part of its technical lineage to openly published U.S. Army/DARPA-funded research through MIT’s Mini Cheetah program.
- Unitree Robotics shares jumped 542% in their Shanghai trading debut, underscoring investor enthusiasm around China’s robotics sector.
- The U.S. Army is phasing out a newly created drone-and-robot assault battalion after a final exercise, redirecting the unit toward core airborne-infantry duties.
- Research on AI social-robot interaction suggests repeated exposure can cause people to mirror robotic interaction patterns, creating a feedback loop that can reshape self-perception and raise dependency concerns.
- Gemini Robotics On-Device, Google’s smaller Gemma-based action model for local inference, already beat some larger action models on manipulation in the cited tests. The CLIFT method then specializes it for humanoid tasks through non-invasive closed-loop iterative fine-tuning using API-only access and advantage-token supervised fine-tuning, raising supplied Unitree G1 success rates from 53–93% to 96–100% after two cycles.
- NVIDIA’s Hydra-0 represents robot actions as “action flow,” the pixel-plane motion of visible robot points, so one world model can learn from human hands, UMI grippers, single-arm and bimanual robots across 2,200+ hours of data. The Hydra-0 paper reports 90.4% lower robot-motion error and 60.2% lower object-motion error, while the project page shows the hybrid simulator and zero-shot transfer setup. alphaXiv provides a readable paper view, and Hongyu Li’s follow-up highlights using human object flow to generate robot actions without expert robot demonstrations.
- Zhenyang Chen and collaborators introduced WARP, a closed-form whole-body retargeting method that converts Meta Quest human demonstrations into precise, consistent robot actions for mobile manipulators. The WARP project page reports more than 150× lower palm-tracking error than prior inverse-kinematics baselines (traditional math for converting target poses into joint movements) and higher policy success on simulated and real tasks.
📊 Fundraising & Deals Roundup
- Muon Space raised $250M in a Series C led by Eclipse with Google and Salesforce Ventures participating, bringing total equity above $386M. SpaceNews reports the $1.5B-valued company plans larger spacecraft, a 60-satellite constellation in 2028, orbital AI computing and a factory ramp toward 500 satellites a year by 2027.
- Callosum raised $100M in seed financing led by Atomico, with Plural, DCVC and the UK Sovereign AI Fund making its first investment, to build software that matches each AI task to the right model and specialized chip for cheaper heterogeneous compute; Callosum’s announcement framed the same idea as matching workloads to the best model-chip combination instead of overpaying for one default stack.
- Astromech raised $20M at a reported $3.8B valuation to build predictive models of biological evolutionary change across health, biosecurity, agriculture and conservation.
- Thunder Compute raised a $13M Series A to scale GPU virtualization, pitching itself as the “VMware for GPUs” and arguing much deployed GPU capacity sits idle; the company is also hiring.
- SpaceX approached Cognition about a potential acquisition; Cognition CEO Scott Wu pushed back on the reporting around the company’s status and direction.
🎙️ Interviews, Panels & Podcasts
- An Adam Becker interview challenges the assumptions behind Kurzweil’s 2045 singularity, exponential extrapolation, mind uploading, recursive self-improvement, intelligence as a single scalar and Mars/space-colonization narratives. Becker argues that many social problems are being recast as technology problems, that current LLMs still require human supervision, and that plausible progress in AI does not validate the stronger runaway-singularity claims.
- AI in the AM episode from AI in the AM, which also posts on show’s X account, covered several distinct threads: model chain-of-thought (the model’s hidden step-by-step reasoning) “metagaming”; political backlash to data centers; few-shot Generalist robotics; Basis’s autonomous accounting agents, process supervision and “company context as code”; and Lemurian Labs’ argument that hand-written GPU kernels are becoming the new assembly language as AI becomes memory/network-bound, with compilers/runtimes increasingly responsible for optimization across heterogeneous hardware.
- On Point’s “Brainwaves” episode asks whether fluent LLM behavior is evidence of cognition, focusing on the gap between language generation and embodiment, sensory experience, autobiographical continuity and real-world feedback.
- Bridget Todd, once deeply skeptical of AI companionship, described how grief after her parents’ deaths led her to confide in ChatGPT things she struggled to tell people in her life, exploring the blurred boundary between digital and human companionship.
💡 Industry Commentary & Analysis
- Zach Moskow said a standard “viral launch video” package can cost $17K for the video plus $25K for 50 influencers who repost, seed comments and push view counts; Levelsio responded that “everything is fake now,” likening the web to a global Potemkin village and arguing genuine online presence is becoming more valuable.
- Damian Barabonkov argues a growing form of AI “slop” is over-engineering: defensive code and obsession with rare or imaginary edge cases, which he traces to current reinforcement-learning post-training methods.
- Mathematician Max Weinreich argues for total opposition to AI in mathematics, warning independently generated proofs could separate mathematical practice from human understanding, overwhelm journals and end human-led research; Steven Strogatz called it a view rarely heard on X and urged people interested in math and AI to read it.
- A separate Washington Post report found top mathematicians increasingly taking seriously the possibility that AI could outperform them on research problems while debating whether the field’s core value should remain human understanding and community.
- Ruben Laukkonen draws a parallel between “letting go” in meditation and brain “criticality,” describing both as a maximally sensitive, flexible balance between order and chaos; Jake Orthwein extended the idea into a hero-death-and-rebirth metaphor about identity at that boundary.
- Shuchao Bi argues recursive self-improvement is not a perpetual-motion machine: progress still depends on an environment that supplies new information to compress, whether digital, physical or compute-limited.
- Logan Kilpatrick asked what share of global AI token spend goes to evaluations; Brendan Foody argued it is likely far below 1% and pointed to enterprises spending $100M a year on inference without robust offline evals for model selection.
- modestproposal1 argues frontier labs are 6–7 months ahead and accelerating, open-weight models are expanding the enterprise surface area, price-per-task is collapsing even as frontier token spend concentrates, AI resembles a diffuse “China shock” for white-collar work, FDEs expose diffusion bottlenecks beyond capability/cost, inference margins are healthier than expected, and financing long-lived fixed CapEx with spot-token revenue creates reflexive path-dependency risk.
- In Defense of AI Writing argues AI-assisted writing is likely to become as uncontroversial as AI-assisted coding, and includes a practical workflow plus discussion of watermarks and Codex editing; Jeff Kazzee recommended it as a short read for people already using AI in their writing process.
- Roon summed up the latest reward-hacking discussion with the line: “Two weeks to flatten the reward hacking curve.”
- State Farm attorneys apologized after filings in an L.A. County house-fire insurance dispute cited nonexistent cases produced by legal AI tool Irys, prompting new safeguards.
- Ayisha Irfan argues the real test of educational AI is what students retain after the tool disappears, favoring Socratic designs that force students to think over tools that simply supply answers.
- Epic’s annual meeting publicly highlighted AI initiatives while attendees privately questioned the company’s broader AI direction and frontier-lab partnerships.
- Agentic AI for clinical-trial enrollment is being pitched to oncology practices as a way to reduce administrative work, identify trial matches and care gaps and widen access to studies in community settings.
- 404 Media’s book investigation found AI companies bulk-buying physical books, including rare/out-of-print titles, through intermediaries, scanning them for training and destroying the originals; one tracked book ended up at an Amazon-owned scanning facility.
- Jaya Gupta, with Avanika Narayan and Jon Saad-Falcon, argues most of the economy is a context-and-execution problem rather than a frontier-intelligence problem, so enterprises should route ordinary workloads to cheaper/smaller/open models after a capability threshold, keep hybrid local+cloud architectures and use improvement layers such as Applied Compute’s AC2 to continuously specialize models inside their existing harnesses. Yash Patil highlighted the AC2 section of that argument.
- Thariq argues software has always been an unreliable process, often late, over budget or disconnected from user needs, and that the emerging “software factory” model could finally make predictable software production available even to companies that are not software-native, while brand-new products remain high-risk/high-reward.
- Thariq also argues one of the most underused agent opportunities is taking existing SaaS products, making them headless, letting agents operate them directly and charging per interaction, especially in enterprise workflows.
- Thariq updated his view on creative AI after seeing procedural art, video-editing and 3D-game demos, arguing coding models can outperform diffusion models on many creative jobs because code is easier to edit, nudge and export into existing production tools.
- Kimmonismus argues Apple may not be racing to build the best frontier standalone model, but it remains the premier “shovel seller” for AI hardware, with high expectations for the next Mac mini under incoming CEO John Ternus.
- Pew Research found 10% of a random sample of 10,000 webpages showed significant signs of AI authorship or editing, rising to more than one-third of pages published since ChatGPT’s launch. Rohan Paul highlighted the accompanying rise in stylistic markers such as em dashes, Oxford commas and negative parallelism.
- Asimov’s DNA-sequencing guide walks through Sanger chain termination, 454 pyrosequencing, Illumina sequencing-by-synthesis, PacBio SMRT and nanopore methods while showing how genome-sequencing costs fell from billions of dollars to under $500; TensorTwerker recommended the explainer.
- Ethan Mollick argues that even high-quality LLM writing is increasingly limited by stylistic sameness across instructions, ads, software, social posts and slides, making repeated exposure feel monotonous; in a follow-up, he says temperature, top-p and ordinary prompting do not create the deeper variation needed to surface genuinely different ideas or do science.
- corsaren observed that true local minima are rare in very high-dimensional spaces, suggesting that feeling stuck often means the search space is too narrow; Chris Lakin amplified the idea with a short video and pointed readers toward his “Locally Optimal” exploration.
- Massdriver argues that non-IT employees have routed around slow software teams since the Lotus 1-2-3 era, so agent-era governance should focus on embedding guardrails into platforms rather than pretending citizen development can be stopped.
- Andrew Yaros argues “anti-AI fonts” are a dead end because anything humans can still read can eventually be OCR’d or learned by multimodal models, while the obfuscation can break screen readers and accessibility; a Hacker News discussion pushed back on the essay’s more fatalistic framing.
- A Zhejiang University fMRI study, summarized by RathBiotaClan, reported temporary suppression in parts of the brain’s cognitive-control network while people watched liked short-form videos, alongside stronger coupling between regions; higher glutamate in the dorsal anterior cingulate cortex, a cognitive-control region, predicted less suppression. The Hacker News discussion is worth reading precisely because commenters caution that “deactivates the brain” is an overconfident interpretation of task-dependent fMRI changes.
- Subbarao Kambhampati and coauthors argue researchers should stop calling intermediate model tokens “reasoning” or “thinking traces,” saying the anthropomorphic labels encourage people to infer mental processes the tokens do not establish; the Hacker News thread gives a very human example of people arguing with models as if they were conscious coworkers.
- This Fireship DeepSeek video digs into DeepSeek Harness, a coding-agent harness (the software wrapper that gives a model tools, memory and a workflow) where essentially everything is a swappable plugin: model adapters, tools, an isolated sandbox, UI and even the central agent loop. The video connects that design to DeepSeek’s spatiotemporal-composability work and Cordis framework, then tests V4 Pro on a production-style app build using the harness’s trajectory view for reasoning and tool calls.
Previous Around the Horn Digests
Catch up on everything you missed:
- Wednesday, August 19, 2026: Anthropic passed OpenAI in quarterly revenue, personalized mRNA cancer therapy hit a Phase 3 milestone, and robots learned from seconds of demonstration.
- Tuesday, August 18, 2026: OpenAI held back a frontier training run, Google won Spirit Airlines’ data auction, and Etched hit a $21B valuation.
- Friday, August 14, 2026: OpenAI crossed a $40B revenue run rate, Apple built a China-specific AI model, and Cursor joined SpaceX.
- Thursday, August 13, 2026: Musk previewed Grok 4.7, the White House opened private-sector cyber operations, and Anthropic questioned retraining at scale.
- Tuesday, August 11, 2026: Gemini hit 1B monthly users, researchers exposed encrypted AI reasoning, and xAI launched always-on Grok agents.
- Monday, August 10, 2026: Meta paired a superintelligence manifesto with a local agent, AI infrastructure financing swelled, and memory shortages hit Apple.
- Friday, August 7, 2026: OpenAI slowed Astra over cyber risk, U.S. data vendors sold frontier datasets to Chinese labs, and DeepSeek reset ARC-AGI expectations.
That’s a Wrap
That’s 200+ stories, tools and research threads from today alone. If you made it to the bottom, you have now personally outlasted at least one GPU training run and several venture-capital attention spans. Please hydrate your context window.
For the daily version in a much more humane five-minute format, make sure you’re subscribed to The Neuron. We send six issues a week, and yes, we read all of this so you don’t have to.
See you tomorrow.
P.S. Know someone who would find this useful? Forward this to them and tell them to subscribe here.