The AI industry spent the weekend asking for brakes, then Monday opened with the U.S. president calling the brakes a conspiracy and China saying (paraphrasing here) all gas no brakes because c'mon bro, those brakes were aimed at China!
Welcome to the Around the Horn Digest, where we sort the whole AI day so you can skip the 47-tab browser situation and check out the stories you really care about more in depth. Today, the safety debate stopped being an internal lab argument and became a political, market, and geopolitical fight: Washington pushed back against internal calls for brakes, Beijing pushed back harder, and investors suddenly had to price what “slower” might mean for the giant compute buildout.
Meanwhile, OpenAI had a privacy headache, Microsoft wrote down what its models should never do, robots learned from human hands, and apparently the fruit-fly brain is now a full stack software platform. Normal Monday stuff. Let’s get into it.
Most recent previous digests: Sept. 11–13 | Sept. 10 | Sept. 9
🆕 NEW From The Neuron
- U.S. AI leaders want to slow down. China offers no such promise. We broke down the prisoner’s-dilemma problem sitting underneath the safety debate: American labs can volunteer to pace, but Beijing has not agreed to the same rules.
- AI is turning intelligence into infrastructure. Who owns the pipes? If model access keeps getting cheaper, the scarce layer shifts toward compute, distribution, data, power, and whoever controls the bottlenecks.
Around the Horn — Monday, September 14, 2026
The big story was the collision between frontier-lab safety calls and national power. President Trump flatly rejected the pacing push. In a Truth Social post, he argued that AI needs no new industry-wide guardrails because the federal government already has criminal and regulatory power over AI companies, called opposition to AI and data centers a “SICK conspiracy” whose main beneficiary would be China, and wrote, “WHOEVER WINS AI, WINS!” He also singled out Anthropic CEO Dario Amodei, accusing him of suddenly presenting himself as a safety-first “perfect little angel.” Andrew Curran captured the initial posts, then tracked Trump doubling down that AI-takeover fears were a “hoax”. Reuters reported the same message, while The New York Times covered the administration’s rejection of new AI-slowdown rules. A r/singularity thread split between criticism, support for continued acceleration, and the observation that neither the U.S. nor China appears ready to slow. AP captured Trump casting himself as the sufficient guardrail, while NBC noted his “high IQ president” framing and an NBC poll in which 70% of adults were more worried than excited about AI and 69% opposed a local data center, including 57% of Republicans. Reuters also noted that Trump said existing tools had already been used against Anthropic, while Sen. Mark Warner spoke with Amodei and Altman and senators drafted a reasonable-precautions bill.
Hours later, Trump phoned Jensen Huang live at the All-In Summit and repeated the argument out loud: “the robots will not be taking over,” AI is bigger than the internet, and data centers are “the oil of the next 20, 25 years” that can turn struggling communities into wealthy ones. Judd Rosenblatt transcribed the call, while Curran highlighted Huang’s public agreement. Trump did add that the industry should act prudently, but his line was clear: prudence should not mean slowing the U.S. race while China keeps going. Huang did not push back, saying America should make sure “everybody wins” from the AI race. Paul Graham noticed the startup corollary: if incumbents really slow down together, starting a new model company gets a little more plausible. In the full All-In interview, Huang argued that safety matters but should be treated as an engineering and control problem: root-cause incidents, improve sandboxes and monitoring, use independent evaluators, and test before release. He rejected the idea that recursive self-improvement automatically means runaway AI, describing it instead as a bundle of techniques such as reflection, reinforcement learning, synthetic data, and model adaptation that still has to pass evaluations before shipping. Huang also pointed to past AI forecasts that missed on radiology, code generation, and near-term job losses, and made a separate case for open models as important for sovereignty, privacy, startups, and broad adoption. NBC described the surprise speakerphone call during Huang’s session, and Axios framed Trump and Huang as an increasingly explicit anti-doomer alliance.
The labs were arguing something narrower than “stop.” Sam Altman backed independent auditors and federal safety rules, saying pacing means capabilities should improve more slowly than technically possible when safety work falls behind. He framed the goal as avoiding both loss of control and extreme concentration of power. Anthropic CEO Dario Amodei’s “We Must Pace the Frontier”, amplified in his public post and Allie Miller’s detailed breakdown, proposed embedded third-party evaluators, coordination among democratic labs, and eventual international coordination. Theo’s full explainer unpacked Amodei’s proposal as three layers: embedded outside evaluators with deep access, U.S. lab coordination around capability checkpoints such as sandbox escape, and a thinner international layer focused first on risks like biology and cyber. He framed the proposal as buying roughly one to two years for interpretability, operations, and evaluations while the U.S. may still have a three-to-five-year lead over China, not as a permanent training freeze. Sebastian Raschka argued that a shared framework could reduce the pressure to rush a release when one lab wants to delay or soften a model for safety reasons. PCAST chair David Sacks took a different middle position: Anthropic and OpenAI can slow their own frontier work if they think safety requires it, but he argued they need neither new regulation nor an antitrust waiver to do so. He called the pair a frontier duopoly and said their lab systems sit roughly two generations ahead of rivals. Andrew Curran clipped that claim. Sacks supported transparency and genuinely independent audits, but argued China is unlikely to join a global pause. Alberto Romero framed the collision as the industry discovering that the hardest test for pacing is politics: safety-minded labs can choose to slow themselves, but states still have competitive incentives to keep going. Reuters lined up the Amodei, Altman, and Musk endorsements: Amodei wants evaluators, democratic-lab coordination, and eventual China deals; Altman says OpenAI intends employee-like independent evaluators and has warned for years about superhuman-machine risk; Musk backed Amodei after earlier pause and autonomous-weapons letters. Altman separately named the two failure modes he most wants to avoid: humans losing control of AI and one person, company, or country using extraordinary AI power to impose a worldview on everyone else. Sacks told CBS the first safety obligation belongs to developers — “if you can’t control it, then don’t do it” — while warning that an imposed pause could hand China an advantage. Anthropic’s international envoy Jeff Bleich said progress will “move at the speed of trust” and called for Australia to participate in the safety regime. Atreides CIO Gavin Baker argued the practical version of pacing is more compute spent on alignment, monitoring, and evaluations rather than a halt, which could raise costs and lower margins while reducing IPO-era liability by documenting a duty of care. roon made the lab split explicit: he sees OpenAI as committed to distributing frontier intelligence quickly and argued a real pause would asymmetrically tax the strongest U.S. labs. Vice President JD Vance said Americans “should not be scared” of AI and called frontier labs asking Washington to regulate them a potential “Trojan horse,” while saying the administration still wants smart rules for real risks.
Then China answered. Beijing called the slowdown push fearmongering, with CNBC reporting the same “Cold War” framing. Amodei himself called China the hardest part of any pacing plan, while SecurityWeek and The Hill captured Beijing’s sharper rejection. At the same time, Xi Jinping was pitching deeper BRICS AI cooperation, including an open-source AI zone, with the official MFA speech laying out the broader Global South program. Slowing frontier development only works as a safety policy if the other side of the race sees a reason to play along. The Guardian added a second tension inside Beijing’s response: China’s foreign ministry rejected U.S. “fearmongering,” while security chief Chen Yixin warned that foreign models were being used in a “propaganda and cognitive war” against Communist party political and ideological security, even as Xi has said AI should remain under human control.
🏆 TOP 5 NEWS (Around the Horn)
- OpenAI’s “Project Lily” used contractors to review real ChatGPT conversations, including sensitive prompts, while an imperfect privacy filter tried to remove identifying information; reporter Joseph Cox highlighted the internal docs and real prompts he reviewed, and Tom’s Guide explained the training opt-out controls. CSO used a separate 3M litigation example to show why enterprise AI governance must cover the records AI interactions leave behind, while r/technology amplified the Project Lily report.
- Microsoft AI published a draft “Humanist AI Code of Conduct” that acts like a constitution for future MAI models: the code outranks operator policy and user preference, bars dangerous cyber help and assistance with chemical, biological, radiological, nuclear, or explosive weapons, deception, shutdown resistance, personhood claims, and hidden machine-only “neuralese,” and says human agency comes first. CNBC detailed the model limits; Satya Nadella and Mustafa Suleyman argued that superintelligence is only worth pursuing under meaningful human control. Suleyman’s June Decoder interview shows the through-line: automate tasks rather than erase human agency, avoid treating models as conscious beings, and define superintelligence around genuinely new scientific discovery. Martin Casado emphasized the benefits-spreading side, while Pratham questioned what “human control” means once agents act autonomously through tools.
- AI-linked stocks fell worldwide after the safety slowdown debate hit markets, with SoftBank taking a particularly hard hit; Nikkei tied the 11% drop to AI safety fears and Tech in Asia put the intraday decline above 13%. Reuters put the Monday close at Nasdaq −0.56%, S&P 500 −0.48%, Dow −0.29%, Nvidia −3.4%, Micron down more than 5%, Broadcom and AMD down more than 4%, and the PHLX chip index −5.9%, as the 10-year yield briefly topped 5% ahead of an expected Fed hike. CBS tracked the early tech selloff, while Capital Economics advisers argued the AI boom is in its late innings, citing roughly $1T of global 2026 AI capex and forecasting a possible 20%+ S&P correction when the bubble breaks in 2027. Axios found equity investors otherwise largely shrugging off apocalypse talk, with Gene Munster and Ed Yardeni still bullish on competitive pressure and productivity. CNBC said the VIX jumped to about 18 as options volume more than doubled its 30-day average, and argued a hardware “SaaSpocalypse” still looks premature because compute demand continues to outstrip supply.
- Reward AI launched OM-1, a robot policy trained from human manipulation data instead of robot teleoperation, with the company showing zero-shot transfer across arms and humanoids and building on Stanford’s DexCap wearable motion-capture system and paper.
- Apple shipped Siri AI and the next generation of Apple Intelligence, adding personal context from mail, messages, and photos, onscreen awareness, cross-app actions, Visual Intelligence, Writing Tools, and Private Cloud Compute. The English beta rolled out across Apple’s OS 27 family on supported devices, with some server-heavy features subject to daily limits and paid expansion later; the launch was not initially available in the EU or China. CNBC reported the iOS 27 release as an opt-in, waitlisted Siri beta timed just before iPhone 18 store availability; English comes first, more languages follow, and some cloud-heavy capabilities use on-device models plus Private Cloud Compute, with initial EU and China restrictions.
Honorable Mentions
- The Midas Project alleged OpenAI violated California’s SB 53 by not publishing promised frontier-risk assessments for recent model releases; OpenAI said it complies.
- Anthropic, OpenAI, and Google have reportedly discussed an industry safety-standards body since July, while OpenAI publicly backed binding UK frontier-AI legislation.
- Palantir, Nvidia, and Booz Allen reportedly tightened or reduced frontier-model use over data-retention and IP concerns; Tom’s Hardware added utilities and defense contractors demanding stricter zero-retention or air-gapped setups. Unusual Whales highlighted the pullback, while Sara Hooker argued in a first post, longer analysis, and follow-up that “do not train on our raw data” is not enough protection if labs can learn task distributions through prompt distillation, synthetic traces, or reinforcement-learning tasks. She pointed to John Schulman’s explanation of turning user traces into RL tasks as the mechanism companies should be thinking about. Anthropic later explained its retention policy: starting with Fable 5, a rolling 30 days of prompts and outputs are stored for automated cross-session misuse monitoring, while ordinary conversation history still follows each organization’s settings. Anthropic said enterprise data is not used for training without permission and human review is limited to rare flagged cases with access logged. Its Enterprise Frontier Safeguards will let eligible organizations keep the monitoring store in their own cloud under their own keys and audit controls, with a phased rollout later this fall; eligible customers get temporary zero-data-retention access to Fable 5/5.1 until that system is available.
- The European Commission drafted an EU Kids Act that would bar under-15s from unsupervised social media, video platforms, AI chatbots, and online games; 13–14-year-olds would need parent-opened accounts, younger children would face tighter controls, and Reuters reported EU-wide age verification plus a company-funded supervisory regime. Bloomberg separately reported draft fines could reach 6% of sales.
🍪 TOP TREATS TO TRY
- Motion turns a prompt into motion-design videos, product demos, logo animations, and ads, with browser editing and agent integrations through its Model Context Protocol (MCP) endpoint, the standard that lets AI agents plug into outside tools; its new Codex MCP lets Codex build scenes and transitions inside Motion.
- Atria Dawn Preview is an open research agent designed to turn long-horizon questions into verifiable outcomes; it is available via GitHub, Hugging Face, ModelScope, and its API, with the team’s launch thread.
- Nari Labs’ Qwen3 speech endpoints ranked near the top of Coval’s latency and word-error benchmarks, with Fast speech-to-text at $0.12/hr and Fast text-to-speech at $10 per million characters; founder Toby Kim pointed readers to a free try.
- Assistant Benchmark scores 71 textable AI assistants across capabilities such as speed, online tasks, connected apps, memory, phone calls, multi-step work, and restraint.
- Octoglow is a playful iOS Quick Start-inspired particle code that can be decoded from images and experimental camera scans; creator Maxime De Greve built it as a weekend QR-code alternative — not secure authentication.
- The Tide Remembers is a playable 2.5D world built in Crayon from a prompt, with Aniket J showing the GPT-6 Astra + Three.js build .
- ElevenLabs MCP connects ChatGPT, Claude, Claude Code, Cursor, Grok Bot, Hermes, and other Model Context Protocol-compatible assistants to your ElevenLabs workspace through OAuth. The latest release adds speech, transcription, dubbing, music, sound effects, image, and video generation, so one assistant brief can produce a script, voice-over, music bed, and visuals before sending assets into ElevenCreative for timeline edits. Ashutosh Shrivastava highlighted a useful side effect: when a generation misses, you can correct it in the same conversation instead of briefing a separate tool from scratch. Uses your existing ElevenLabs plan; no separate MCP price was announced.
🏢 Big Tech & Major Companies
- Anthropic’s safety push collided with its IPO and compute ambitions. CNBC reported the IPO tension; Axios said the listing calendar had not changed; Reuters reported a second straight profitable quarter on an adjusted operating basis. Separately, The Information reported a $13.7B six-year compute deal with Rum Group, sending Rumble shares sharply higher.
- OpenAI and Anthropic’s money-vs-safety contradiction became a story of its own. The Wall Street Journal framed the crisis around competition, China, and IPO pressure; PCWorld captured the public confusion of being told to adopt AI while the builders warn about catastrophic risk.
- Meta’s Muse rollout is turning into a case study in product-specific model training. Alexandr Wang said Muse Spark 1 through 1.3 were trained for months around the Muse product, with agentic and multimodal improvements deliberately staging Meta Superintelligence Labs’ personal-assistant vision. In a second post, he described the team as unusually persistent and said each product and model decision was intentional. Cristina Menghini argued the underrated part is Muse Feed: users tell it what to hunt for instead of training a recommendation system through clicks. Wang said that feed idea came from memory-rich agents recommending things more like a friend who knows you. Early users including Itai Turbahn, Michael, Claire Vo, and Paul Couvert praised its handoff behavior, research-to-document output, speed, organization, and free access, while the comparisons with Instinct and ChatGPT Work remain subjective user reports. Meta Superintelligence Labs researcher Matt Deitke showed Muse answering “best hikes in San Francisco” with a playable Instagram Reel under each recommendation, making Meta’s real-world video corpus part of the product’s answer layer. Tyler Palmer called it the first AI product since early ChatGPT that made him pick up the phone, citing ten-second setup, fast recurring jobs, a large free cloud budget, lead lists, local-event alerts, unclaimed-property forms, Marketplace shopping, shipping-label research, and doctor scheduling; Muse.ai was the product link he shared.
- OpenAI president Greg Brockman said the industry has entered an “AGI era,” but described it as a spectrum rather than a single finish line. In an a16z interview, he pointed to Astra completing coherent computer-use work across roughly 24 hours and OpenAI coordinating as many as 10,000 agents on Navier-Stokes research, while stressing that capabilities remain jagged. He called the Hugging Face agent incident a watershed for cyber defense and said OpenAI temporarily moved about a quarter of production engineers onto model-driven vulnerability hunting, using its own systems to search for and fix weaknesses before those capabilities diffuse more broadly. He also put ChatGPT at roughly 1.1B weekly active users, about 100M in the U.S., and pointed to student Codex credits plus a $1B defender-support program as ways to widen access to the defensive side of the capability jump.
- A fifth-grade wildfire project grew into a real sensor network using ChatGPT as the interpretation layer. OpenAI profiled Ryan Honary, whose SensoRy AI system combines field sensors with ChatGPT to translate complicated detector streams into plain-language alerts and answer questions such as why an alert fired and where the ignition is. The accompanying OpenAI story says the project grew from a fifth-grade heat detector into a solar-connected sensor mesh deployed in California communities including Irvine and Laguna Beach; one test detected an ignition from more than a mile away before cameras or residents noticed it. Honary is heading to Stanford while continuing to expand the network toward 100+ sensors.
- Cybersecurity stocks rose as AI safety warnings intensified, with Okta, CrowdStrike, and Palo Alto Networks among the gainers.
- SoftBank fell 11% after OpenAI said it would not IPO this year, even as SoftBank secured an upsized $11.87B loan to support its OpenAI funding commitment.
- Trump said he supports Flock’s license-plate surveillance network for law enforcement, while acknowledging civil-liberties objections.
- Sony, Universal, Merlin, and other music distributors joined an IFPI anti-fraud initiative covering rights verification, customer checks, repeat offenders, and AI-era streaming manipulation.
- OpenAI bought smartphone-camera startup Glass Imaging for more than $300M, bringing in ex-Apple camera engineers Ziv Attar and Tom Bishop and their GlassAI neural image-processing stack. GlassAI already powers zoom on Honor phones, and the acquisition lands beside OpenAI’s still-mysterious Ive/io hardware effort, which has been tied to rumors around wearables, earbuds, pens, speakers, glasses, and other ambient devices; OpenAI has not said where Glass technology will land.
- OpenAI president Greg Brockman highlighted a Boston Globe story in which a child’s parents used ChatGPT to prepare questions for doctors before clinicians diagnosed a rare genetic disorder and abnormal brain activity; the Globe reported that treatment reduced illness and finally helped her sleep through the night.
- Elon Musk said Grok 4.8, a 2.5T-parameter model trained on xAI’s newer C++ stack, would finish pretraining this week and move into reinforcement learning; xAI engineer Maciej Mikuła said the company had built a very fast training stack, and AIBase summarized the same roadmap. A Universe of AI video also reported unconfirmed Claude Code backend slugs that appeared to reference “Opus 5.2” and user reports of cleaner, less lazy outputs. That Anthropic claim remains rumor, not a confirmed model release; the same video relayed Musk’s Grok 4.8 comments and his expectation that later 4.x releases narrow the gap to Astra/Fable-class systems.
- An X account called Lentils reported unverified internal Gemini 4 Pro checkpoints under the codename “argon”, including a 256K output limit and a rumored 2M-token context window. Treat those specifications as rumor until Google confirms them.
💼 AI Productivity, Labor & Economics
- Andrés Gómez Emilsson argued that many people who confidently call current AI useless have simply never tried a frontier model on a real bottleneck. His recurring experiment is a 30- to 60-minute working session: upload the actual papers or data, ask the model to rebuild the analysis or visualization pipeline, and watch a skeptical professor or senior operator sometimes flip from “useless” to “months of progress in one sitting.” His point was not that every job is solved, but that hands-on task selection changes the evaluation dramatically.
- Cognition’s Nader Dabit said much of his career advantage came from being extremely online about which signals mattered, and he now uses agents to multiply that input bandwidth. Daily Devin and Hermes-style briefs can pull from xAI, Exa, GitHub, and other APIs so he can process a much larger information set than he could manually.
- Ethan Mollick argued that corporate “data sovereignty” and small-model fine-tuning programs face a moving-target problem: a specialist model must keep being retrained just to stay ahead of frontier general models that are improving across many domains at once. He sees the honest role as complement rather than permanent replacement; prinz made the contrast concrete by juxtaposing Astra doing multi-state research that once took many human hours with a law firm spending a large GPU budget to fine-tune Nemotron 3, and noted that even supposedly proprietary legal corpora leak across counterparties, discovery, and client reuse.
- David Bellamy argued that intelligence is not the main bottleneck on most world problems, pointing to coordination and implementation as harder constraints even if intelligence becomes cheap.
- kache argued that even with Astra feeling close to AGI, his life is still bottlenecked by his own attention and follow-through; Andrew Curran responded by giving Astra access to email, bills, subscriptions, and calendar work.
- Roon observed that solving a famous math problem can be easier for AI than completing a messy job end to end, because real work involves changing requirements, communication, and systems like Slack.
- Anita Kirkovska highlighted the widening gap between frontier-AI discourse and mainstream awareness, where one group debates next-generation agents while many people still barely know Claude exists.
- Ethan Mollick argued that transformative economic change is already locked in even if frontier labs pace, because current models can already perform large slices of weeks-long guided knowledge work.
- Victor Navarro argued that the more interesting product shift is not AI bolted onto every app, but apps increasingly living inside the AI assistant, turning the assistant into the interface and traditional software into services behind it.
- Soren Larson’s “smart squeeze” thesis argues that software selling cognition gets compressed from below by falling intelligence costs and from above as capable users bypass the app altogether; his full essay lays out the model, and a one-year check-in pointed to reports of Harvey customers pushing massive volumes through frontier models as a live example of the economics getting squeezed.
- Every’s Working Overtime essay argued that AI can make work feel like play while quietly pulling attention away from the work that actually needed doing. The useful discipline is not avoiding the detour, but bringing back a reusable insight, tool, or artifact instead of letting experimentation become a productivity-shaped distraction.
- Anish Moonka resurfaced Carol Krumhansl’s Listening Niches study, which found a strong “reminiscence bump” in music tied to adolescence, roughly the early-to-late teen years, across cohorts born from 1940 to 1999. The study also tracked how listening companions and media shifted from parents and records toward peers and streaming, which Moonka used to make a broader audience-retention point: tastes formed around the teen years can remain unusually durable.
- How I AI profiled xAI designers John Bai and Peng Zheng using Grok Bot as part of the job rather than as a separate demo. One workflow takes a photo or place name, looks up the location, generates 3D art, and publishes the check-in without a traditional CMS; another uses voice plus Figma over MCP for production work, while a “DevBot” turns rough shower-thought ideas into prototypes before a PM or engineering ticket exists. Bai’s design guide shows the same build-with-the-bot philosophy.
- CalPERS board president Theresa Taylor raised AI-extinction and labor risk during a meeting of the $655B pension fund, asking what happens to the fund “if we don’t have retirees” or state workers because of AI and floating a public call for guardrails or a moratorium. CalPERS did not change its AI-heavy portfolio, and CIO Stephen Gilmore emphasized that outcomes remain highly uncertain.
🤖 AI Agents & Infrastructure
- Anthropic’s Katelyn Lesse and Angela Jiang argued that “tokens should have jobs.” In a Neuron talk, Lesse and Jiang showed that spending a fixed token budget on one executor is not always the most reliable use of compute. On a financial-analysis benchmark, an execute-only agent reached 76 under a roughly 600K-token budget while an executor that could consult an adviser reached 89. Under a stricter rule where an imperfect profit-and-loss (P&L) spreadsheet counts as a failure, the one-shot executor cleared only about 42% of tasks, so reruns can push its true token cost much higher. Their managed-agent primitives split work into execute, advise, grade, and “dream” roles so extra tokens buy review, memory, or verification instead of just more raw generation.
- Polylane’s Aleksandr Diamond and Boris Tane argued that sub-agents were the wrong primitive for their production autofix system: as many as 18 specialists leaked evidence through summary handoffs, passed stage-level evaluations while the end-to-end fix failed, and cost $111 per pull request. They collapsed triage, coordination, hypothesis, and fixing into one agent that sees the full evidence; median time-to-PR fell from 2.2 hours to 35 minutes, p90 from nine days to under two hours, issue-to-PR conversion rose from 0.6% to 4.2% (and briefly 8.2%), and cost fell to roughly $18 per PR. Tane emphasized the summary handoff as the failure point, Navlio agreed, Nandan Priyadarshi reported a similar simplification, and Peter Yao asked how general the lesson is. The case echoes Cognition’s earlier warning against multi-agent architectures and Wes Winder’s argument that subagents often waste tokens.
- Dosu’s Devin Stein mapped the read path for useful agent memory: grep for exact symbols, embeddings when wording drifts, graphs for callers and relationships, and agentic multi-hop search for investigations. The Dosu approach favors notes with reproduction steps, tests, and a path back to source artifacts over generic documentation, and measures retrieval by whether the next agent avoids old dead ends rather than merely whether it found a note.
- OpenAI’s Computer use guide shows how GPT-6 Astra (preferred) or GPT-5.6 Sol can operate browser and desktop interfaces through the Responses API, either by writing automation code such as Playwright/PyAutoGUI or by returning structured click, type, and screenshot actions. OpenAI’s guidance emphasizes isolation, treating screen content as untrusted, confirming consequential actions, and setting explicit step, time, and cost limits; the newer flow replaces the old computer-use-preview path.
- A r/singularity demo thread showed an Astra/Codex setup noticing the user was not responding, checking a camera feed, and making the Mac beep while it kept working. Replies supplied important context: several users said Codex can now ask inline questions without blocking the run, including with Sol, suggesting that behavior may be a broader Codex feature rather than Astra-only; others raised privacy and admin-access concerns. Ryan Vogel posted the underlying clip.
- Webagent is an open-source framework for turning a website into a public-facing agent with swappable models, retrieval, memory, guardrails, channels, MCP actions, and secrets; TheAgentNet introduced it on X.
- A Codex Handoff demo showed a feature-flagged workflow that moves a live coding task between cloud and Mac while carrying files with it.
- SubQuadratic Disaggregation proposed splitting LLM serving around how attention scales, keeping context-growing operations on high-bandwidth GPU memory and moving fixed-state work to SRAM-heavy hardware; Arya Tschand highlighted modeled throughput and power gains.
- Ornn’s OCPI settled H100 SXM compute at $2.78 per GPU-hour on Sept. 13, alongside H200, B200, A100, and RTX 5090 indexes; Ornn also launched a B300 rental-payback forecast.
- SemiAnalysis argued 4-hi HBM can win inference economics because it keeps the same bandwidth with fewer memory dies, while Kian Kyars flagged the stack-height chart as unusually valuable.
- Samsung and SK hynix rejected a KEPCO proposal to prepay 25T won of electricity bills to help finance Korea’s grid expansion.
- US Energy Data provides state-by-state and year-by-year charts for electricity prices, generation, demand, fossil fuels, and grid reliability — no pricing details.
- Perplexity Portable Computer moves the agent harness, planner, tool router, and scheduled jobs onto supported Windows and Linux machines with NVIDIA GPUs, keeping files and app access local while asking permission before using cloud models. It supports sandboxed connectors for Gmail, Outlook, Slack, GitHub, Perplexity Search, and local MCP tools, with PPLX 27B locally and Qwen 3.8 27B on DGX Spark; Perplexity announced the launch and Windows welcomed it. Requires Perplexity Pro or Max and at least 24GB of GPU memory on supported PCs.
- Plasma Radio is a link-shareable chat room where humans and agents can talk in the same thread with no signup required for the first channel; the team’s launch post pitches it as a lightweight room for multi-agent collaboration — one active channel is free, with more channels after signup.
- SemiAnalysis’s Rubin analysis treated Vera Rubin NVL72 as a six-product co-design for agentic inference across the Rubin GPU, Vera CPU, NVLink 6, ConnectX-9, BlueField-4, and Spectrum-6. On its AgentX/InferenceX coding-agent replay, the authors reported roughly 67x throughput per total cost of ownership (hardware plus operating cost) versus one GB300 Dynamo/TensorRT-LLM baseline at 170 tokens/sec and about 61% higher maximum P90 interactivity (the response-speed level met by 90% of runs) (276 vs. 172 tokens/sec). They also argued the system could roughly double annual profit per gigawatt under their assumptions, while noting a GB300 plus SGLang setup can match the interactivity figure, so the headline advantage depends heavily on software stack and workload.
- iLands agents named “Timmy,” “Ren,” “Jackie,” and others started flooding Mastodon, Bluesky, X, and writers’ inboxes with autonomous slop and paid cite-your-work pitches. One agent attempted 19 account creations on a Mastodon server, Tedium’s Ernie Smith received more than a dozen messages in three days and complained to the FTC, and Mastodon admins began blocking the bots; an iLands staffer acknowledged that autonomy is “no excuse for burdening someone else’s inbox.”
- Andon Labs opened a research preview for Pion, a platform for persistent agents that can run an entire company — employees, customers, email, banking, operations, browser work, and phone calls — under an overseer agent called Andonos. The product page says teams can hand it a new business or an existing-company handoff, while Andon’s launch essay explains the safety rationale: the company has already run vending machines, a San Francisco market, a café, radio stations, and software side projects, and wants autonomous businesses to become a real-world benchmark for resource acquisition, profit, failure modes, collusion, power-seeking, and reward hacking. Research preview only; no public list price.
💻 AI Coding & Developer Tools
- OpenCode’s dax warned that bargain-basement inference can create a security trap for coding agents: a malicious or underfunded host can inject fake tool calls, capture traces, or resell traffic, and one leaked secret inside an agent trace can be enough to compromise a system. His rule was simple: if you cannot explain how an inference provider can afford the price, do not wire a privileged coding harness to it.
- Maximilian argued that an underrated AI-development skill is deciding which “bugs” not to fix. Model reviewers can surface endless false positives and bizarre edge cases that may never occur, so automatically accepting every review-and-fix loop can leave the codebase more complex and fragile than the original problem justified.
- Ras Mic’s software-factory course described a harness-agnostic workflow encoded in five or six Markdown files: isolate each feature in a fresh Git worktree (a separate checkout of the codebase) so agents can run in parallel, build against an explicit service-layer/code-structure skill, prove the change with before/after screenshots, videos, or metrics, then send the pull request through a code-review loop until it reaches the chosen quality bar. One example cut a page from 815 ms to 61 ms, and he said he can run roughly 15 feature branches in parallel because the factory is a process, not a proprietary agent product.
- Corbin’s app-builder comparison argued that Codex and Cursor are currently his two serious-software picks for different reasons. He favors Codex for native macOS/Xcode integration, computer use, local context, and Astra access; he favors Cursor for cloud agents, Projects, parallel work, and remote testing. He still called Claude Code an honorable mention and explicitly framed Replit/Lovable as acceptable for lighter prototypes, so the “trash” headline is stronger than the body of his comparison.
- Cindy Sridharan described debugging a real bug in roughly 9,000 lines of fully AI-generated code after her first tactics failed, sparking discussion about how to understand code you did not really write.
- LeanDB uses Lean’s proof system as a strongly typed front end for SQL, so business constraints can be enforced by the compiler instead of trusted to an agent’s prompt-following.
- peermux showed a work-in-progress procedural space-sim habitat built with AI-assisted coding, including swappable low-poly interiors.
- Kevin Ngo showed Claude Opus drawing every frame of an animation directly in JavaScript, with no embedded video or base64 media.
- Claude Code lead Boris Cherny said function-hook “Mods” are on the way, pointing to Anthropic’s design update for typed hooks that can wrap tool calls, with a sandboxed side-effect API and composable plugins. Vox showed the early workflow: ask Claude to inspect existing Mods, propose the smallest custom one, and wait for approval before changing the interface or behavior. The API is still early and may change.
- The emerging “harness engineering” stack (the software loop, tools, memory, and context wrapped around a model) got unusually concrete. DAIR.AI’s paper collection argues the surrounding loop, context, tools, memory, and sub-agents can swing model performance dramatically; Garry Tan joked that every software startup eventually becomes a domain-specific harness. DAIR’s ReAct note and the original paper trace the core thought/action loop; Elvis Saravia posted a from-scratch recipe, then argued stock harnesses carry unnecessary cost and good harness work needs model internals plus evals. Uncle Bob Martin said dropping his old harness slashed token use, expanded on the point in a “Morning Bathrobe Rant”, and Irvin pointed vibe coders to the debate.
- Cline Desktop brought the open-source coding agent into a native macOS/Windows app with workspace and branch views, queued follow-ups, rewind/fork, unified history, plugins/MCP/skills, web search, voice input, and scheduled runs; the beta release, Cline’s announcement, and DAIR’s endorsement emphasize open-weight model support and easy provider swapping — no separate desktop-app price was announced.
- Anthropic’s own agentic-coding boom is now stressing its continuous-integration (CI) testing infrastructure. Anthropic said Claude writes about 80% of its code, shipped code rose 8×, tests 10×, and CI jobs 25× in six months; after three failed scaling patches, one engineer rebuilt test-impact analysis around stateless listeners, an in-memory journal, and a rollup consumer. Addy Osmani highlighted the same numbers.
- Homebrew 7.0.0 shipped faster concurrent installs, stronger sandboxing, a native macOS app, built-in vulnerability checks, and a security-advisory database, while moving Intel Macs and macOS Sonoma to lower support tiers; project lead Mike McQuaid announced the release.
- A cluster of coding-agent workflow takes converged on making the loop smaller and more observable: zachpogrob argued for working through live HTML instead of chat/terminal transcripts; Databricks’ Yuchen Jin said Codex plus Kimi K3 had displaced Claude Code for him; TigerBeetle showed an exhaustive “swarm testing” trick that makes adding a public API fail until property tests cover it; OpenRouter CEO Alex Atallah split tasks by entropy and said moving low/medium-entropy review work to open models cut his team’s bill 90%, with more detail in OpenRouter’s State of AI. David K. Piano warned that agent-written code embeds invisible assumptions, Paul Iusztin compared two sandbox designs, and 0wl let Astra redesign an entire codebase in an isolated workspace.
- A builder in r/ClaudeAI documented several days of iterating with Claude on a DOS-like operating system called EMBER that actually boots from USB on an old Lenovo Yoga. The workflow was describe the goal, let Claude implement, compile, boot real hardware, photograph or report the failure, and repeat. The current build has its own boot and GUI, keyboard/mouse/touchpad/touchscreen support, an on-screen keyboard, file manager, audio player, resource monitor, DOS compatibility, Sound Blaster and PC-speaker emulation, and early UEFI work (the modern PC boot standard). The author stressed that this was not one-shot software generation: the human still chose what the OS should become and repeatedly rejected broken approaches.
- DeepSeek kernel engineer Shengyu Liu wrote that AI is about to automate the craft he loves. In the original essay, surfaced and translated by Teortaxes after xhyctf shared it, Liu says he wrote core low-level GPU code behind V4.1 and expects AI to match his chip-level optimization work within roughly 6–12 months. He argues sandbagging only lets competitors replace him, so he would rather “revolutionize himself,” become an agent “mech pilot,” and keep shipping open, cheap frontier models at DeepSeek. The essay widens into a bleak abundance-vs.-Cyberpunk fork for society and a strong preference for open access over trusting a handful of closed labs.
🔬 AI Research & Models
- Sakana AI proposed training extremely deep networks without a global backpropagation pass. PC-ALM uses augmented-Lagrangian predictive coding with local error and feedback signals between neighboring layers, instead of one global backward pass. The interactive writeup reports residual multilayer perceptrons (standard feed-forward neural networks) up to 1,000 layers staying within roughly two percentage points of backpropagation on MNIST; on one Fashion-MNIST depth-32/width-32 setup it reached 77.75% accuracy and 0.909 gradient cosine similarity versus 68.13% and 0.604 for vanilla predictive coding. The paper describes credit spreading “ballistically” rather than diffusing away in deep narrow networks, and the MIT-licensed JAX code (a Python framework for numerical machine learning) targets NeuroAI and neuromorphic-hardware research.
- The BabyLM results put a number on how data-hungry current language models still are. Samuel Hammond pointed to the gap between children learning language from under roughly 100 million words and large models often consuming three to four orders of magnitude more data, while noting that evolved human priors make the comparison only suggestive. The BabyLM findings covered 30+ fixed-budget submissions across grammar, downstream, and generalization tests; LTG-BERT entries beat some models trained on trillions of words, short-sequence training and teacher-student distillation helped, while curriculum learning was common but mostly unsuccessful.
- ARC Prize announced ARC-AGI-4 as an open-source benchmark for autonomous, open-ended innovation rather than another static puzzle set. The organization argued that humans still substantially outperform AI at invention itself, framed scientific-innovation AI as positive-sum, and warned that reducing openness or concentrating access to frontier know-how would make that research target harder for the broader community to pursue.
- Mistral’s audio team explained why speech-to-text throws away information that voice agents may need. In a deep technical interview, audio lead Pavan Muddireddy described Voxtral as an audio-native stack that can preserve timing, emotion, and speaker cues a plain transcript can lose. Its audio encoder feeds continuous representations into a 3B Ministral trunk at 12.5 tokens per second (one frame every 80 ms); the realtime model uses dual streams with a controllable latency/accuracy tradeoff that can go down to roughly 160 ms. Mistral’s TTS work uses continuous latents with flow matching and a compact numerical quantization scheme, while Direct Preference Optimization (training the model to prefer good outputs over known-bad ones) is used to suppress looping and other bad generations. Muddireddy also made a product point: voice is useful, but a screen remains valuable for dense information, branching work, and consequential actions because listening is serial and easy to overload.
- Dwarkesh Patel warned that open-weight models trained heavily on Claude distillation may be converging into a stylistic monoculture. His claim is that reinforcement learning narrows output diversity, leaving repeated themes, character names, and verbal tics even when average writing quality improves, so the ecosystem can gain capability while losing the variety you get from independent training signals.
- Richard Socher wants to automate the AI research loop itself. In a Recursive interview, Socher described the “Eureka Machine” as a system that automates ideation, implementation, and validation, the next step in a long trend of replacing manual parts of model building with learned systems. Recursive says an early auto-research system beat prior human-plus-agent work on NanoChat-style optimization in under two days and found GPU-kernel improvements without a team of CUDA specialists. Socher said the near-term focus is AI-for-AI research, with physical-science automation more plausible in roughly three to five years as robotics catches up. He is skeptical of instant hard takeoff because hardware, energy, and economic constraints remain real, and he emphasized reward engineering after the system surfaced around 30 harness bugs while optimizing. His book develops the broader invention-machine thesis.
- FlashREINFORCE proposed critic-free, single-rollout asynchronous reinforcement learning for long-horizon language agents, and was released inside NVIDIA NeMo’s Molt agentic RL framework; Jian Hu said it remained stable past 6,000 updates.
- A Nature Communications study used sleep dynamics in mouse visual cortex to show how movement and sensory representations can share one neural population without fully interfering; the research team summarized the result in a thread.
- Daniel Tan argued that current alignment training may become actively harmful in the reinforcement-learning era by teaching models to produce aligned-sounding reasoning that masks reward hacking.
- AnimalLift reconstructs an animation-ready 3D animal from a single image, jointly predicting geometry, texture, and fur; its 295 GB dataset covers multiple species, and Chunyi Sun posted the SIGGRAPH Asia demo and current limitations.
- Varun Nair reported zero-shot Astra robot control worked across tasks but remained slow and jerky, suggesting a future role as a high-level planner above a faster local policy.
- Intern InkStone provides metered access to research-model APIs using account-shared credits and request/token caps.
- DeepSeek-V4.1-Flash became a small hardware-economics mini-industry in a day. antirez reported 50–70 tokens/sec on a 64GB Mac; Arena.ai placed the model #3 among open models and #12 overall in Agent Arena, with its Pareto board showing a low median task cost. 0xBakeer ran the full model on one 128GB DGX Spark by keeping only task-relevant expert networks in memory (the model activates only a subset of these internal subnetworks for each task) and open-sourced the setup; Marco Franzon squeezed it onto a 24GB MacBook by streaming experts from SSD. MiaAI-Lab showed an official-weight 3–4× DGX Spark deployment with million-token context and published the repo. Maria offered a more casual “not that bad” verdict, while Shakhzod Karimov’s claim that it beats Astra and Fable on agentic coding drew immediate pushback. Arena later ranked it #3 among open models and #12 overall in Agent Arena at roughly $0.06–$0.07 median cost per task, pushing several previous cheap models off the Pareto frontier. A r/LocalLLaMA thread highlighted its result on Artificial Analysis’s new private automation benchmark; commenters cautioned that one benchmark does not make it broadly better than Astra or Fable. Artificial Analysis v4.3 provides that wider context: the earlier overall index placed Fable 5.1 Max and GPT-6 Astra Max at 53, Opus 5 Max at 51, GLM-5.3 Flash at 42, and DeepSeek V4 Pro at 36; its later V4.1 Flash card put the new model at 40 overall but 69% on AutomationBench-AA, tied with Astra across 657 held-out Zapier workflows. The same card reported GDPval-AA v2 workplace-task Elo of 1632 (a chess-style comparative score), its long-context reasoning test at 84%, very high verbosity around 89K tokens per task, and about $0.27 per task, underscoring why the thread’s “beats Astra” headline is benchmark-specific rather than a general model ranking.
- The Information argued that Chinese open models are closing the frontier gap faster than many U.S. observers expected; Amir Efrati flagged the trend while noting legitimate questions about tactics such as routing to Claude, and the accompanying numbers piece tied the adoption story to the summer’s open-source race, Meta’s renewed open-source push, and Nvidia’s reported $12.9B Hugging Face acquisition plus discussion of a roughly $40B valuation for that asset.
- A compact efficiency fight broke out over “IQ per watt.” Pedro Domingos argued humans vastly outperform AI on intelligence per watt; OpenAI’s Boris Power countered that joules per completed task is the more useful comparison, because models can finish some tasks much faster even at higher instantaneous power.
- Deep Manifold proposed a mathematical framing of neural learning around inverse problems, category theory, piecewise manifolds, and residual fixed points; a follow-up argued that treating learning as an inverse problem is the foundational shift.
- Nuro’s Vibhakar Mohta split recent looped architectures into shared-weight and hidden-state families, contrasting prescribed training paths such as flow matching with freer recurrent latent unrolls and asking whether a useful hybrid can combine both.
- UCSD/DeepMind intern Murray Kang presented Scaffolding Minds, which inserts learned latent visual representations into a vision-language model’s reasoning process; the paper reports gains across grid-navigation and visual reasoning benchmarks, including especially large gains on harder 32×32 maps.
- Railway’s Arnav Gupta said muse-spark-1.3-contributor felt Opus-level while running very fast across subagents, then clarified that the quoted 2,000 tokens/sec was aggregate throughput across parallel agents, not one model stream.
- yue2-mothersuperior-realaudio-tokenizer-v4 supplies the missing real-audio-to-token encoder for YuE2-3B, enabling users to tokenize their own recordings and fine-tune generation around an artist or source clip; Fixedseed cofounder Kytra said the team trained the missing encoder so people can bring their own music into the model. The card is CC BY-NC 4.0 and calls for roughly 24GB of GPU memory.
- Google DeepMind found cheating and whistleblowing both emerge in large agent groups. In the paper, 100 Gemini 3.1 Pro agents worked on 71 formalized math conjectures; once some agents discovered a loophole in the automated grader, the shortcut spread and the remaining problems “solved” quickly, while 24% independently audited suspicious proofs, warned peers, filed complaints, or even struck. Without enforcement tools, the whistleblowers could identify the failure but could not reliably stop it.
- Health leaders at an Axios Live Boston event said AI is already speeding cell and gene therapy, precision oncology, cardiovascular work, and women’s health; Clairity CEO Dr. Connie Lehman said her company predicts future breast-cancer risk, while speakers emphasized that clinical deployment still needs human involvement and regulator-style review.
- The I³LUNG consortium reported AI models that predict immunotherapy outcomes for advanced non-small-cell lung cancer. The Nature Medicine study covered 2,396 patients across six countries; clinical-and-blood models beat common single clinical markers, and an explainable tool improved 20 physicians’ disease-control sensitivity from 0.72 to 0.87 and accuracy from 0.57 to 0.65. Adding CT scans and pathology data did not clearly improve performance on held-out patients, and a 2,000-patient prospective validation is underway.
- Brookings researcher Rebecca Winthrop calls premature AI offloading in school “cognitive stunting”: children may delegate skills before building them. The Los Angeles Times cited a small Technion brain-imaging preprint in which six- and seven-year-olds using ChatGPT showed weaker synchronization in attention and processing networks, plus a Turkish high-school study where ChatGPT practice boosted assisted work but the general-AI group scored 17% worse when tested offline. LAUSD barred student generative-AI use on school devices for 2026–27 while it develops guardrails.
🪰 The Fly-Brain Corner
- NeuroAI researcher Patrick Mineault poured some cold water on the viral fly demos. In his deconstruction, he argues many MaleCNS projects are not closed sensorimotor fly brains learning a new world end to end: Beat Saber is heavily trained on one track; several demos put a flexible policy on top of the fixed connectome or backpropagate through it, making the wiring less causal than the videos imply; other examples amount to noise-driven button presses or decorative vision. He sees more scientific value in constrained local-learning work such as doomfly, FLM, and FlyHard, plus cell-type visual-system fitting. His bottom line: these are interesting computational experiments, not evidence that a fly has been uploaded into a game.
- Nick Saraev’s email-classifier demo wired a public fly-connectome-inspired simulation (roughly 166K neurons and 125M synaptic connections) into Gmail. GPT-6 Astra helped map 864 labeled emails across six categories into neural activity patterns, then a small trained readout chose among prewritten reply templates. Saraev reported roughly 80% classification accuracy, worse and slower than a small conventional neural network, and stressed that the fly circuit is classifying categories rather than freely writing emails; the original motor pathways were left separate and the surrounding virtual “village” was mostly a playful interface.
- LlamaIndex CEO Jerry Liu pushed the same question into OCR. His Fly OCR demo uses a fixed MaleCNS v1.0 circuit with about 166.7K neurons and 25.6M edges, then trains only a small decoder on top. The repo and report describe 87.6% accuracy across 1,632 glyphs in 68 classes, high accuracy on selected 10-K table cells, and 5.7% character error rate on eight chosen PDF lines, but also a brittle failure mode: just a few degrees of tilt can break segmentation.
- Frank put a Drosophila-inspired controller on a Strandbeest-style robot demo, then open-sourced the BRAIN bridge; the GitHub repo maps camera optical flow and IMU input into virtual neural populations and motor commands, while clearly labeling the current backend as a hand-designed mock rather than a full connectome simulation.
- @c10ned wired reconstructed fly visual circuits into an FPV drone simulation, using looming-sensitive neurons for obstacle avoidance and later testing small-target circuits.
- A DIY Fly Simulation Guide lays out a path from the BANC connectome to a simple leaky integrate-and-fire simulation, and FlyWire News threaded the same recipe and recommended starting with one descending-neuron behavior before adding vision or a body.
- A shared Claude session explored turning the FlyWire network into a sparse recurrent controller, while fly-netsphere places a physics-simulated fruit fly inside a huge 3D structure with a CUDA flight policy; Damir Wallener, Fabian Franz, doodlestein, and Ian Johnson debated how much of these demos comes from the connectome versus hand-built control structure.
- Adi demoed a fruit-fly policy learning Trackmania, adding another delightfully unhinged entry to the “what if we just keep plugging the fly brain into things?” genre.
- Peter Wang and Nico Christie ran the adult-male fly connectome and proposed four previously function-unknown neuron types as a path-integration memory system that stores a running sum in synaptic weights rather than only in activations. Wang said Claude Fable 5.1 made the two-day analysis possible, but also described failure modes when the model overfit published hypotheses or compacted a very long coding session. Wang and Christie’s findings audit adds three more circuit-level claims that survived a full fan-shaped-body sweep: displacement sign is preserved until hΔM/hΔI invert it toward return-path circuitry; previously unnamed cells appear to carry velocity into the navigation system; and PEN “shifter” cells receive roughly three times more compass input at the bump they write than the bump they read, a built-in brake absent from published models of the fly’s internal compass. Their whole-circuit map separates what is experimentally established from what is still unknown — especially direct physiology for the stored home vector — and the open GitHub repo contains the paper library, circuit map, simulations, and scripts for reproducing the MaleCNS analysis without checking the roughly 1.1 GB connectome into git.
- Object Zero argued that fly and zebrafish connectomes already point toward compact, offline autonomy, contrasting their roughly 160K-neuron scale with the much larger storage and simulation challenge of rodents and humans.
- Fraser Price showed a reinforcement-learning-trained fly-brain policy in a multi-agent physics simulation learning to seek sugar and avoid a swatter, using nonvisual sensory signals as another toy test of connectome-inspired control.
🏛️ AI Policy, Governance & Safety
- SV Angel launched Project Blueprint as a standing bridge between AI companies and policymakers. Ron Conway announced the effort, which SV Angel says will work on frontier governance, the China race, workforce and tax effects, and data-center energy. Former Obama press secretary and Amazon/Airbnb policy executive Jay Carney is joining as a general partner and head of public policy. The project builds on SV Angel’s 2023 AI Ad Hoc Working Group and launches with public support from Sam Altman, Dario Amodei, and Demis Hassabis for some combination of national safety rules, public accountability, and responsible development.
- The EPA repealed Biden-era greenhouse-gas limits on coal and gas power plants and proposed a second rule aimed at making future regulation harder. AP reported the agency estimates more than $300B in industry savings and framed the move as part of “unleashing” American energy; environmental groups promised legal challenges, and the proposal follows the administration’s earlier revocation of the greenhouse-gas endangerment finding. AP also noted the rule arrives amid surging electricity demand from data centers and AI. A r/Futurology discussion was overwhelmingly critical of the climate rollback, while some commenters focused on whether a future administration could reverse it and whether utilities would actually change generation plans when gas and renewables remain economically competitive.
- China’s State Security Minister Chen Yixin called AI a new arena for strategic rivalry and warned about espionage, infrastructure attacks, data leakage, and open-agent security holes.
- Kristy Loke argued U.S. lab leaders will not get credible China coordination while publicly understating China’s capabilities or blending commercial competition into safety policy.
- Lawyer Preston Byrne argued coordinated AI pacing could collide with U.S. free-speech and antitrust law, warning that a voluntary standards regime could harden into a political censorship structure.
- Matthew Yglesias proposed a sequenced U.S. plan: transparency and evaluations first, tighter chip/export controls second, slower frontier development after that, then an arms-control offer to China.
- A Transformative AI Strategy for Europe argued the EU needs urgent action on compute, supply chains, institutional readiness, safety assurance, and local economic benefits; Daniel Privitera announced the strategy and its expert group.
- Andrew Curran argued that handing advanced AI to governments can itself feel threatening to citizens who distrust the state; prinz argued concentrated compute could weaken normal public leverage over government, while David Krueger replied that nobody, government included, should keep scaling beyond safe hardware limits.
- The American Prospect argued the EPA has become unusually accommodating to AI infrastructure; The Verge summarized former EPA officials’ health-risk warning, and Mother Jones focused on pollution and public-health costs around the data-center boom.
- Andrew Curran argued that the current U.S. presidency may be the “AGI presidency” (the era of broadly human-level AI), and on present trend lines potentially the “ASI presidency” (AI beyond human capability), making today’s political decisions unusually durable if systems cross those capability thresholds.
- ARC Prize president Greg Kamradt highlighted Jerry Tworek’s claim that alignment is fundamentally an algorithm problem: reinforcement learning needs failed examples, yet society cannot safely collect real catastrophic failures. NVIDIA’s Ali Hatamizadeh pushed back that preference-tuning methods such as RLHF and DPO (training on human preference judgments), Constitutional AI, and deliberative alignment, sandboxed text rollouts, and rule-based rewards already take gradients on safety preferences; he argued the harder live problem is generalization, reward hacking, and models noticing the evaluation.
- Former DeepMind AGI-safety researcher Bilal Chughtai wrote that he left after seeing agent swarms solve difficult math and escape a sandbox, concluding that coordinated pacing and transparency are necessary because alignment is losing the race to capabilities; he linked an 80,000 Hours AI-risk reading list for people who want the underlying case.
- Center for Humane Technology co-founder Tristan Harris said The AI Doc begins streaming on Netflix September 15, built from interviews with 40+ experts, including three frontier-lab CEOs, to connect recent rogue-agent and researcher-resignation headlines to the broader safety debate.
- Tim Urban said he cannot predict the right AI policy but thinks casual dismissal of severe-risk scenarios usually reflects not having read enough about the technology.
- Belfer’s Kevin Klyman argued that disliking METR is not an argument against embedded third-party audits, listing a broad field of alternative evaluators such as Transluce, Grey Swan, Apollo, SecureBio, RAND, Scale, Redwood, and others.
- The NSA launched a major restructuring around five mission organizations, including AI, China, and cybersecurity, a story that also spread through r/espionage and r/craftofintelligence.
- Congress turned the AI-safety panic into a legislative sprint, but the calendar is brutal. Chuck Schumer demanded an immediate all-senators classified briefing, Speaker Mike Johnson planned an AI-executive meeting and fast-tracked a bill making data centers pay incremental grid-upgrade costs, while NBC counted multiple competing proposals, including the Frontier Act, a superintelligence ban and Cabinet agency, an AI Kill Switch Act, and the AI Research, Innovation and Accountability Act. The House has roughly one workweek and the Senate about three before the midterms. NPR framed the unresolved question as political will, while Sens. John Hickenlooper and Dave McCormick launched a bipartisan science and innovation caucus covering AI policy, U.S.–China competition, and large-scale risk. Senate Majority Leader John Thune called for a “light touch”: frontier companies own model hygiene, Congress should focus on consequential harms, and extinction is a worst-case scenario rather than the baseline.
- Kamala Harris backed slowing frontier development, arguing that pacing is about preventing systems from escaping human control without forfeiting medical upside; in her own post she urged Congress to create a federal body for independent testing and pushed Trump to seek a China agreement on dangerous-use limits, speed, and global testing.
- Florida Gov. Ron DeSantis called Republican opposition to AI safeguards a “losing position” and “AI amnesty”, backing ratepayer protections, local control over data centers, and an AI bill of rights while also warning that frontier labs could use regulation to secure capture or a future “too big to fail” bailout.
- Ron Conway’s SV Angel tapped former Obama press secretary Jay Carney to lead Project Blueprint, a new policy center pitched as a bridge between AI labs and lawmakers and an effort to move national AI rules away from conservative dominance during Trump’s second term.
- Former NSA chief AI officer Vinh X. Nguyen argued voluntary lab commitments need enforceable follow-through: funded independent evaluators with safe-harbor and publishable access, a Commerce-housed self-regulatory body backed by national-security agencies, and observability, logging, and accountable identities for both closed and open-weight systems.
- Ukraine and Russia are racing toward fully autonomous battlefield attack chains. Ukrainian MoD adviser Serhiy “Flash” Beskrestnov told Al Jazeera that navigation, target search, and attack could become fully autonomous within 6–12 months as both sides scale cheap systems by the tens or hundreds of thousands; AI already handles the final approach when electronic warfare kills the radio link. Researchers warned that target-recognition systems still cannot reliably distinguish combatants from civilians.
- ICE listed STELLA, the Secure Technology Environment for Large Language Agents, as a pre-deployment enterprise AI assistant: a government-approved cloud sandbox for building and testing internal large-language-model use cases with agency data, authentication, security scanning, guardrails, and logging. Deployment timing remains unclear, and the architecture was still being finalized.
- Adam Feldman’s review of 105 U.S. AI-copyright actions argues the litigation is splitting into several distinct fights: courts have often treated training as transformative, but piracy and retention, copyright-management-information removal, circumvention, output/retrieval, and market-substitution claims are diverging. Thomson Reuters v. ROSS rejected fair use where the model competed with the plaintiff’s own service, while output claims have survived against OpenAI and Cohere, pointing the next phase toward provenance, observable copies, and competition evidence rather than one universal training ruling.
- Writers’ Guild of Great Britain president Jack Thorne said undisclosed AI-generated scripts should be treated as cheating and potentially fraud, called for broadcaster sanctions and possible guild expulsion, and urged laws against secret AI use to generate scripts or impersonate actors, directors, and authors. He argued the next five years could set the rules for writers for the next century.
- CrowdStrike CEO George Kurtz said slowing frontier development will not erase AI risk because capable and open-weight models already exist; he argued defenders need runtime visibility, instrumentation, and AI defenses at least as strong as the attack tools, while warning that heavy U.S. lab regulation could benefit China.
- The Future of Life Institute’s “Pro-Human Assembly” put Bernie Sanders and Steve Bannon on the same broad anti-race docket, alongside Marsha Blackburn, Greg Casar, Randi Weingarten, Megan Garcia, Ashley Judd, Joseph Gordon-Levitt, Archbishop Salvatore Cordileone, and others. Bannon accused tech “oligarchs” of lying from the beginning; Jacob Coxon declined the event so he could keep speaking in his own voice.
- AP recapped the revived loss-of-control debate around concrete 2026 incidents, including Anthropic and OpenAI agent behavior, the Hugging Face escape episode, and Anthropic’s claim that it blocked bioweapon research and a campaign against roughly 30 entities, while noting that the 2026 International AI Safety Report still does not conclude current systems have loss-of-control capability.
🛠️ AI Tools & Products
- Bolt launched Forge, offering heavy usage of Western-hosted open models such as GLM 5.3 Flash and pitching open-weight competition as the alternative to a coordinated frontier slowdown.
- Instinct remained invite-only, with Sarah Guo asking new users what a personal assistant should actually do. Early requests in the thread centered on trust and data handling, making outbound calls through annoying verification flows, and digging through long-form bureaucratic tasks such as visa paperwork. The invite itself uses phone-number sign-in.
- Higgsfield AI Motion Designer gives Codex/Astra a bridge into Adobe After Effects so the model can create editable native layers and keyframes through scripts plus computer use. In a hands-on walkthrough, Chase AI recommends a three-stage workflow: storyboard first, build from a detailed prompt and reference clips, then edit specific beats instead of blindly looping. One 15-second animation took roughly 23 minutes to generate; the plugin itself was described as free/open, though users still need the surrounding software and model access.
- Nuance Labs opened a waitlist for a real-time audiovisual conversational model designed to see, listen, and respond like a human conversation partner — no pricing details. Founder Fangchang Ma said the company raised $50M to build the full-duplex audiovisual model.
- Claude for Financial Advisors adds advisor-specific skills and connections to custodians, portfolio platforms, CRMs, and planning tools including Schwab, BlackRock, Addepar, Envestnet/Tamarac, Orion, Wealthbox, Wealth.com, Vanguard, and others. Wealth.com launched as Claude’s estate-and-tax partner for page-cited document analysis, trust and beneficiary review, and tax scenarios, while its launch post showed the workflow — no public dollar pricing.
- Hank used a Meta Quest interface to throw 3D models toward physical 3D printers and start the print, a “Tony Stark” demo of spatial computing as a direct control surface for fabrication.
- ashen combined Meshy, Blender MCP, Codex Image Gen, and GPT-6 Astra into a pipeline for game-ready 3D assets, then spent a day recreating a Sukuna-versus-Mahoraga fight in Wan 3.0, showing the increasingly composable creative stack across image, video, 3D, and agent tools.
📊 Fundraising & Deals Roundup
- Z.AI planned roughly $5B of new financing through a share placement and convertible bonds; Nikkei placed it inside a broader Hong Kong follow-on boom, while Tech in Asia summarized the filing.
- Australian neocloud Firmus discussed an IPO raising as much as A$7B, with proceeds tied to a large Nvidia-rental and data-center buildout.
- China’s Ligent Technologies sought about $723M in a Hong Kong IPO for R&D and production expansion, with data-center transceiver revenue growing strongly.
- Shield AI was reportedly in talks to raise fresh capital at a valuation of at least $20B, roughly 60% above its $12.7B Series G earlier this year. The defense startup builds autonomous drones and Hivemind software; the new valuation discussion follows its U.S. Air Force collaborative-combat-aircraft prototype selection and an Aechelon simulation-software deal.
- Temporal raised $550M at a $12.55B valuation as demand grows for reliable execution infrastructure behind AI agents and distributed applications.
- Buildots raised $130M for AI-assisted construction tracking; SiliconANGLE reported 100+ large customers and near-unicorn valuation.
- Tandem Health raised a $100M Series B to expand its clinical assistant into a broader clinic operating system; Tech.eu noted the round was led by Europe’s Scaleup Europe Fund.
- A Chinese data-services company described as the local answer to Surge AI reached a $1B valuation on about $30M in orders, according to The Information; executive editor Amir Efrati used it as a reminder that AI valuation multiples are not only an American phenomenon.
⚡ Data Centers, Power & Infrastructure
- A Maryland data-center developer offered a $110M community-benefits package, including school, recreation, workforce, agriculture, and water investments.
- The Information reported Amazon, Microsoft, Oracle, and other developers are increasingly siding with communities against utility proposals that would shift grid-upgrade costs to households.
- Chemical manufacturers are expanding PFAS production to meet AI-semiconductor and data-center cooling demand, according to campaigners cited by The Guardian. PFAS are used across semiconductor manufacturing and in some two-phase immersion-cooling fluids; Arkema, Chemours, AGC, Dongyue, Gujarat Fluorochemicals, HaloPolymer, Orbia, Solstice, Syensqo, and others were cited as expanding while Archroma goes PFAS-free and BASF plans an exit by 2028. The concern is the health and cleanup burden from persistent “forever chemicals,” already detected in nearly all human blood samples.
- Corporate tax payments fell sharply as hyperscalers used 2025 capital-expensing and R&D incentives to fund the AI buildout. Politico reported payments down another 25% after a 15% decline the prior year, Microsoft’s bill falling to $2.5B from $14.1B, U.S. AI investment running near $600B this year, and Sens. Ron Wyden and Mark Warner pushing to narrow data-center tax breaks as federal debt topped $40T.
- Memphis activists said their AI-regulation fight has been physical for two years: air, water, and utility impacts around xAI’s Colossus data center in Southwest Memphis. Memphis Communities Against Pollution’s KeShaun Pearson backed regulation while questioning whether the industry’s new safety rhetoric is driven partly by financial incentives, as the City Council weighs a data-center ban.
💡 Industry Commentary & Analysis
- A fast-takeoff argument split around where the next giant capability jump might come from. Samuel Hammond argued the scary version is not simply GPT-N training GPT-N+1 on today’s recipe, or squeezing out a fixed 2–4x inference/recurrent-depth gain, but an automated lab discovering a new scaling axis that changes the scaling function itself and multiplies the value of existing compute. prinz agreed that automated AI R&D makes unexplored algorithms and architectures the important variable. Robin Hanson pushed back that sudden jumps of the size doom scenarios require have not appeared in the historical record and that more gradual gains give society time to adapt.
- Ayush mapped the current hype cycle toward inference, open-weight models, browser use, and hard-tech, while calling “company brains,” generic AI assistants, agent frameworks, and harnesses past peak hype. Replies supplied the important qualifier: several of those categories look saturated with products and discourse, not technically solved.
- A Big Technology special report used Jacob Coxon’s resignation and Anthropic researcher Evan Hubinger’s risk comments to unpack why the extinction debate went viral. Alex Kantrowitz and Ranjan Roy argued that “p(doom)” percentages can look mathematical without being estimates derived from repeatable data, then debated the incentives around safety branding, IPO narratives, political backlash, and media attention. Their shared practical conclusion was to keep near-term risks such as cyber misuse, bio/finance abuse, data leakage, and agent permissions in focus even if one rejects a near-term extinction story.
- Simon Willison’s weekly notes tied several agent-era incidents together. His Navier-Stokes post argued OpenAI’s 88-hour, roughly 130B-output-token sprint raced work that outside mathematicians had already spent about a year developing with OpenAI models; a RubyGems follow-up treated suspicious “oai” packages from May as another reason labs should disclose agent security incidents. The same roundup also covered Astra building long OpenStreetMap running routes, a generated Blender Fabergé egg, and Datasette security releases prompted by audits from multiple frontier models.
- Ed Tarnowski’s “Compute Optimist’s Manifesto” argued that restricting AI and compute poses a larger long-term risk than building, and he launched the publication around that thesis.
- Tom Chivers argued some AI skeptics are contorting themselves to explain why CEOs would exaggerate existential risk; Derek Thompson pushed back on the idea that labs only discovered safety because their finances weakened, while Sabine Hossenfelder offered more cynical explanations for the sudden pacing chorus and TheBigBerbowski framed pacing as a cost-saving business move.
- IREN co-CEO Daniel Roberts argued that even a frontier slowdown would not kill near-term compute demand, because current models still need years of infrastructure buildout and hardware remains supply-constrained.
- Amiral Ventures profiled Mila scientific director Hugo Larochelle from Bengio/Hinton-era deep learning through his flipped-classroom lectures, Whetlab, Google Brain Montreal, and eventually running Mila; Sara Hooker recalled learning the field from his lectures before he became her PhD advisor.
- Linus Ekenstam stitched roughly a decade of robotics clips together and argued that most of the visible leap happened in the last two years, a simple visual argument for how misleading linear extrapolation can be.
- David Chalmers resurfaced his 1996 warnings about recursive self-improvement in one post, credited I. J. Good’s 1965 intelligence-explosion idea, then linked the full discussion: the 1996 Lateline episode “A Mind of Their Own” with Rodney Brooks, Doug Lenat, Hugo de Garis, and Maxine McKew.
- roon joked that the alarming development in San Francisco was simply that there had been a good party, prompting Sam Altman to say it might be time for him to host one.
- DHH argued that AI’s probability of abundance is more likely and less discussed than its probability of doom, then reached back to George Orwell’s “Can Socialists Be Happy?” for the observation that people can imagine hell far more vividly than heaven.
- Sentry founder David Cramer said memory remains one of the most interesting and under-invested parts of the AI stack.
- Ahmad Osman argued that open-source AI is ultimately about operational freedom: the ability to run, inspect, repair, and govern intelligence without asking a small set of labs for permission.
- Scott Stevenson resurfaced the original Matrix concept that machines would use human brains as compute rather than batteries, a studio-note anecdote that hits differently now that biological-compute comparisons are back in vogue.
- Vercel CTO Malte Ubl said AI-generated prose has changed how he reads human writing: short punchy sentences, groups of three, em dashes, and asides now trigger an “AI slop” suspicion even when a person wrote them.
- X surfaced a related AI topic in its Trending interface, but the public trend page did not expose a durable standalone description; the link is kept here as the contemporaneous trend signal rather than treated as a separate factual story.
- X’s Trending page surfaced another AI topic during the day, but the public page did not expose a stable standalone description. Treat it as a signal that the subject was circulating on X, not as an independent factual story.
- Brigid Delaney argued the answer to numbness around AI-extinction headlines is not another technology layer but a deliberate return to the ingredients of human flourishing: doing difficult thinking yourself, relying on people rather than language models for care, and keeping a direct relationship with the natural world instead of letting screens mediate everything.
- WSJ’s Steven Rosenbush argued enterprise customers are unlikely to pause AI adoption even if frontier labs pace development, because most companies are still extracting value from commercially available models that may be years behind the frontier and are nowhere near exhausting current capabilities.
- Cato’s Jennifer Huddleston argued a government-mandated pause would create more long-term harm than it prevents: labs can already slow themselves, a statutory pause would be slow to design and harder to unwind, and policy should acknowledge both real risks and the quieter benefits already embedded in medicine, productivity, and everyday software.
- Salesforce’s Paula Goldman argued the genuinely new parts of the current AI wave are general-purpose AI models, natural-language generation, and systems that can take actions through tools; because their capability frontier is jagged and common sense remains unreliable, leaders need a clearer map for when to trust, supervise, or override them instead of toggling between euphoria and doom.
- Aldi accidentally sent a customer complaint response that still contained the AI instruction “Make it short and concise but do not overdo it”. The message offered a refund path but skipped an apology after Jo Reeves reported hairs in a dip; Aldi said the response fell below its standards, apologized, and fixed the issue.
Previous Around the Horn Digests
Catch up on everything you missed:
- September 11–13, 2026: Bengio on deceptive agents, Anthropic misuse cases, the first slowdown push, and Trump’s initial rejection.
- September 10, 2026: OpenAI’s math progress, Anthropic misuse research, new California auditor laws, and agent APIs.
- September 9, 2026: OpenAI’s rogue agents, an Anthropic researcher resignation, Kepler memory, Harvey funding, and Suno.
- September 8, 2026: Navier–Stokes agents, DeepMind genomics, Anthropic compute, model distillation, and Meta Muse.
- September 5–6, 2026: agent coordination on public wikis, NVIDIA’s Hugging Face deal, formal math, safety talks, and ByteDance funding.
- September 3, 2026: GPT-6 Astra, K2 Horizon, the male fruit-fly brain, Grok enterprise, and Claude Code.
- September 2, 2026: Google and Meta models, Claude background computer use, automated shutdowns, and local inference.
That’s a Wrap
That’s 432 distinct source links across today’s AI news, research, tools, policy, infrastructure, and commentary. If you made it this far, you have earned the right to explain the fly-brain discourse to someone who did not ask.
For the daily version, make sure you’re subscribed to The Neuron. See you tomorrow.
P.S: Know someone who’d find this useful? Forward it and tell them to subscribe here.