Frontier models are now beating licensed accountants on carefully scoped month-end work, while AI labs keep shrinking models, caches, and even the amount of human supervision they need. Friday is apparently efficiency day.
Welcome, humans.
Today’s pile had one theme hiding underneath a lot of very different stories: AI keeps getting cheaper to run, easier to specialize, and better at work that used to require an expert in the loop. That showed up in accounting, Stratego, local models, coding agents, and even VFX.
The uncomfortable part is that the benchmark curves are moving faster than the job descriptions. A spreadsheet task does not equal a career, but “the model now clears the test your juniors take” is still the kind of sentence managers remember. Let’s get into it.
🆕 NEW From The Neuron
- OpenAI Agent Security: The New Rules for Containing AI Agents: why agent containment now looks less like one sandbox and more like layered permissions, monitoring, isolation, and incident response.
- Investigators Found a New Problem With OpenAI’s Rogue Agents: Missing Evidence: the investigation problem is not only what an agent did, but whether the evidence survives across outside services.
- AI Skill of the Day Digest — September 2026, Part 2: 13 practical workflows across ChatGPT, Claude, Gemini, Jev, Codex, and agents.
Around the Horn — Friday, October 2, 2026
Mercor tested 12 licensed CPAs on four simplified month-end close tasks. The humans averaged roughly 37% of rubric criteria. Frontier models that trailed that range last year now cluster near the top, with Claude Opus 5 near 100% on the study’s chart.
The important caveat is the task design. These were well-specified file-and-math jobs, not messy client conversations, judgment calls, or full end-to-end closes. Aden Barton’s thread makes the same point: the result says a lot about structured junior work, not that “accounting is solved.”
Still, the shape of the change is hard to ignore. Matt Stockton called out how fast the curve flipped, while Sheel Mohnot compared the moment to legal AI after GPT-4: agents do the checklist-heavy first pass, humans review the exceptions. That is a much more useful mental model than “AI replaces accountants,” and probably a more disruptive one too.
🏆 TOP 5 NEWS (Around the Horn)
- Ataraxos beat the most decorated Stratego player 15–1–4 over 20 games using self-play plus search under hidden information, then generalized the same approach across other imperfect-information games; code and weights are on GitHub.
- OpenAI Foundation named 163 nonprofits across 35 states and D.C. to split $50M in unrestricted People-First AI Fund grants, with no requirement that recipients use OpenAI products; the Foundation also shared the grantee list on X.
- DeepSeek is reportedly building software for Huawei AI chips, including six open-source tools and a TileLang adaptation for Ascend 950, as Chinese labs keep reducing their dependence on NVIDIA’s software stack.
- SpaceX’s AI unit reportedly accelerated data-center construction with above-market labor, offsite prefabrication, and industrial robots. The Information also reported summer talks to lease compute to Microsoft, with secondary accounts describing Anthropic and Google commitments, a target near 420,000 NVIDIA processors in November, and a December contract around $1.1B a month. The Microsoft piece was talks, not a signed deal.
- PewDiePie was reportedly banned twice by OpenAI while trying to distill Sol using Sol-generated seed data; a longer recap of his path into local AI turned the episode into a wider argument about who gets to learn from whose model outputs.
Honorable Mentions
- Ben Affleck said InterPositive fine-tunes open video models on production-owned footage so crews keep both the dailies and the learning; the full Bloomberg Live interview followed Netflix’s reported $587M acquisition of the 16-person post-production company. Thomas Wolf joked that Affleck had become the timeline’s fine-tuning instructor, but the underlying detail is real: the shop spent eight months shooting its own stage data and only fine-tunes the last cinematic layer for each production.
- Cerebras said disaggregating inference, sending the compute-heavy prefill stage to partner accelerators and keeping bandwidth-heavy token generation on wafer-scale systems, raised throughput 5× in early tests with the same wafer count; Cerebras summarized the result on X.
- Matthew Green argued that sandboxing alone cannot contain useful agents because useful agents need data and network access; Vicki Boykis highlighted the piece as a model for more grounded AI-safety discussion.
- Sam Altman reportedly named two early humanoid-robot jobs: building more OpenAI robots and constructing OpenAI data centers, echoing the recursive automation loop he sketched in The Gentle Singularity.
🍪 TOP TREATS TO TRY
- OpenRouter's Model Router Benchmarks compare seven routing systems on quality, speed, and cost across six benchmarks, with a default Router Index weighted 60% quality, 20% time per task, and 20% cost that readers can reweight. The launch post and technical write-up separate blends, selectors, and switchers and warn that routing can rebuild prompt caches, add latency, and misjudge complexity from the prompt alone; no separate pricing is listed.
- ChatGPT Finances is rolling out in the U.S. with account connections through Plaid and Experian for questions about subscriptions, duplicate charges, rising bills, budgets, credit-score drivers, debt payoff, emergency funds, investments, and large purchases; OpenAI's feature list also shows voice-based planning around a job change. Ethan Bloch says U.S. access spans Free, Go, Plus, and Pro with international access still in progress, while the public ChatGPT homepage does not advertise the feature yet.
- Lightpanda 1.0 took its from-scratch Zig browser for machines out of beta with roughly 1.74M passing Web Platform Test subtests, CORS on by default, private-network blocks, stricter cookies, and output as HTML, Markdown, or a semantic tree. The launch thread says the project crossed 10,000 commits, while the 1.0 release supports Puppeteer, Playwright, Selenium, Chrome DevTools Protocol, and a native MCP server. No pricing details.
- Cloudflare Clef is an Apache-2.0 decision model built from Qwen3.8-27B with a joint schema head that scores allowed answers without generating prose; the smaller Clef-flash targets lower latency. Cloudflare's technical post says Clef beats Jev on several decision benchmarks, while MiaAI highlighted the early speed and accuracy gains. No end-user price is listed.
- Kev 1.0 is Jared Palmer's Apache-2.0 family of Jev-compatible decision models that you can train and serve yourself, from 0.8B to 27B. The GitHub repo includes fine-tuning tools, the weights collection is public, and Palmer's launch says the 27B model reaches 75.7% generalization in his harness with pooled calibration error of 0.019. Weights are free; you pay for your own compute.
- SpaceXAI's experimental TypeScript SDK wraps text, voice, image, and video APIs plus real-time X search, web search, code execution, files, batch, collections search, image generation, remote MCP, and tool search in one server-side client. The Apache-2.0 repo and npm package are public; interfaces may still change before 1.0. The SDK is free, while API usage requires a SpaceXAI key.
- Prime Inference is Prime Intellect's serverless and reserved serving layer for frontier open models across multiple data centers with failover. The launch says its stack serves about 600B tokens a day, while GLM-5.3 on GB200 NVL72 targets 100+ end-to-end tokens per second per user through prefill/decode disaggregation and compressed KV transfer. Its OpenAI-compatible API root is live, but there is no public price list.
- Runware Serverless lets you deploy a Python handler or container on autoscaling GPUs that can scale to zero, with stable APIs, logs, and metrics. The launch post advertises a $0.63/GPU-hour reserved floor; pay-as-you-go starts at $1.60/GPU-hour for L40S and rises by GPU class, billed by the second while a worker runs or stays warm.
- Arceus is a law firm that combines licensed attorneys with agents for contract review, negotiation, and drafting, with flat prices of $500, $1,000, and $2,000 for its main workflows. Mac Liu says the company launched with $17M led by Greycroft; Arceus claims 3–5 hour average turnaround, an eight-hour on-demand guarantee, and 70%+ lower customer legal spend versus traditional firms.
- Reka's Inverse Dynamics Model turns short video clips into low-level camera and movement commands so world models can learn from unlabeled footage. Reka says its flow model, trained only on game footage, reached 85.9% on forward/turn recognition in real video; the Apache-2.0 weights are public.
- Every Agent is a Slack agent for 5–50 person startups and agencies that works inside threads across codebases, PRs, Drive, Gmail, Notion, Figma, GitHub, Ramp, PostHog, HubSpot, and Airtable. Every says beta includes $15 in credits per person, capped at $500 per workspace; after that it is $30/user/month or $24/user/month billed annually, plus model tokens at 0% markup.
- Pruna's P-Video-2-Pro Cost tier sits on DesignArena's preference-versus-price Pareto frontier for both text-to-video and image-to-video, starting at $0.01/second for 480p and $0.025/second for 768p. The docs, API dashboard, and P-Bench are public.
- QuiverAI Arrow rebuilds a sketch or raster into structured, editable vector geometry instead of tracing it into thousands of points; the Quiver workspace is live. No pricing details were provided.
- Inception's Voice Playground lets you talk to Mercury Voice inside a live agent pipeline with tool calls, a per-turn latency breakdown, and a head-to-head against a baseline. Inception claims more than 2× lower latency than GPT-6 Luna and under 300ms median end-to-end latency in a dental-receptionist demo; enterprise access is via sales@inceptionlabs.ai.
- Vals AI's Web Search Index holds the model and agent harness fixed and swaps only the search tool across professional finance and legal tasks. Vals says agents score 2.9% on legal and 7.4% on finance with no search versus roughly 30–50% with search, with Exa, Keenable, Parallel, and Tavily included. No pricing details.
- Choir is an Apache-2.0 protocol for distributed multi-agent autoformalization: a human orchestrator breaks textbook-scale work into GitHub tasks, independent contributors use their own agents, and a deterministic gate checks PRs before human review. The paper covers Lean 4/Mathlib, Isabelle, and Rocq, while Melanie Weber points to a demo formalizing 285 theorems from Yufei Zhao's Probabilistic Methods in Combinatorics. Free to try.
- text-to-cad gives coding agents skills to create and edit CAD from text or images, export STEP/STL/3MF/GLB, pull off-the-shelf parts, write drawings and robot URDF/SDF files, check manufacturability, and slice to G-code or a Bambu job. The project is MIT-licensed and free to try.
- Hermes Desktop now shows a one-click MCP catalog flow inside Capabilities → Connectors, including DeepWiki for asking questions about public GitHub repos. The same install works from the terminal, and the Hermes docs cover the catalog workflow.
- NVIDIA’s 64GB DGX Spark ships October 23 from Acer, ASUS, Dell, Gigabyte, HP, and MSI starting at $4,999. It uses the same GB10 Grace Blackwell chip and ConnectX-7 networking as the 128GB box, targets models up to 100B parameters locally, and two units can pool to 128GB for models up to 200B. MiaAI highlighted NVIDIA’s claim of twice the memory bandwidth and up to 1.7× speed on a Qwen 3.8 27B test.
- Underdog is Sigil Wen’s private personal assistant built around a 27B Qwen 3.8 model and a model-specific Husky engine that he says runs up to 4.5× faster than Apple MLX on the same weights. His Underdog’s Law manifesto argues that today’s frontier intelligence reaches consumer devices in about six months; the product is pitched as free, ad-free, with personal data staying off an Underdog server, and iOS, Windows, Linux, Android, NVIDIA, Omarchy, and a coming Dogpark in the roadmap.
- Cua Spaces gives agents their own desktop on your Mac, hardware you own, or your cloud. Teleport can move an app’s tabs, profile, and sign-ins into a Space after approval, while Cua Driver lets agents operate macOS, Linux, and Windows without stealing focus and the Spaces app is source-available. Cua’s launch says connected-computer Spaces are free, hosted relay is free in early access, and cloud compute is billed by your provider.
- AstaBrief 8B turns a research question plus retrieved literature into a cited report in one pass. Ai2 says the Qwen3-8B-based model was trained with supervised fine-tuning and preference optimization, scored 87.0 versus 77.3 for base Qwen3-8B on ScholarQA-CS2, and powers Asta Fast mode at about 51 seconds per report versus 178.5 seconds for its Claude-powered Thinking mode. The example repo and technical write-up are public; weights use Apache 2.0.
- Whistle is a 16.9MB on-device speech model from Cactus that transcribes seven languages on the same CPU engine as Needle, with 11.1ms first-token latency and 1,319 tokens/second on an M4 Pro. Cactus says it mostly beats Whisper base while being 9× smaller and 6× faster, supports keyword biasing, timestamps, silence handling, 2–4 bit quantization, and 17 deployment targets including browsers and tvOS; no audio leaves the device.
- Rapid-MLX is a local Apple Silicon inference server and Mac app with OpenAI- and Anthropic-compatible endpoints for Claude Code, Codex CLI, Aider, Hermes, and DeepSeek Harness. The Apache-2.0 repo includes 27 tool-call parsers, continuous batching, quantized cache support, and multimodal serving; v0.15.4 added TensorFold acceleration paths for Qwen3.8 and GLM-5.3 Flash.
- TensorFold is Ash Hart’s exact-decoding engine for Apple Silicon and modern NVIDIA GPUs: draft tokens are accepted only when they match serial decoding under the same settings. Volatile Markets showed Nemotron 3.5 Lightning jumping from 178 to 700 tokens/second on an M5 MacBook Pro with byte-identical output.
- MiaAI’s dual-DGX recipe, announced here, serves GLM-5.3 Flash EXL3 across two DGX Sparks with a 1M default context, four concurrent streams, and about 60 tokens/second on one prose stream or 108 across four. Mia is replacing the quantization with a 4-bit TensorFold build that keeps the same 176GB footprint and speed while claiming lower KL divergence and fewer confident reasoning mistakes; weights are public, with another update promised around the new quant.
- Muse Gadgets lets you wire Muse to ESP32 boards, Raspberry Pi devices, displays, buttons, microphones, speakers, sensors, actuators, and Home Assistant using open-source firmware and Linux SDKs. Nat Friedman says Meta also manufactured 5,000 Muse Home Link USB-C devices, free to U.S. Muse subscribers while supplies last.
- Gemma 4’s Developer Agent Competition asks teams to post-train Gemma 4 31B into an offline coding agent that navigates a repository and drafts fixes on consumer hardware, with SWE-bench-style pass/fail grading and a 12-hour patch limit. Google Research says the pool is $100,000, split between the agent track and an optional paper track; entry closes November 25 and final submissions are due December 2.
- Suno Speech takes an idea, poem, or text plus a voice and musical style, then returns spoken audio over an original score for dramatic readings, meditations, voice notes, pep talks, or bedtime stories. It entered beta October 1 after a smaller test.
- Cloudflare K2 is a serverless ordered event log built on R2 object storage, with independent producers and consumers, fan-out, batched reads, ack/nack leases, and long retention without running a Kafka broker cluster. Public beta is available to Workers Paid customers with 10GB and 30MB/s per stream and no beta bill; Cloudflare lists post-beta pricing at $0.04/GB produced or consumed plus $0.02/GB/month retained.
- Gauth launched an Unlimited Digital Canvas tutor where chapters sit side by side, board notes follow narration, students can zoom from concept maps to formulas, pause for answers tied to a step, run Quick Checks, export PDFs, and resume from a shared link. Gauth.com also offers Gauth Atlas visual lessons, rapid STEM walkthroughs, and access to verified tutors; the Product Hunt listing says the broader platform has passed 200M downloads since 2020. The launch was hunted by Aleksandar Blazhev, whose profile also surfaces Scholé, Scarlett, and Questflow.
- Teachoo is a free homework helper that takes a typed problem or photo and works toward the answer one question at a time, with guided questions, feedback, larger hints, and step explanations. Its own site covers math, science, history, and writing and emphasizes learning by solving rather than simply revealing the answer.
- Discourse is Erik Torenberg’s Cosign-graph forum between broadcast social and private group chat: about 5,000 whitelisted people can post directly, everyone else goes through a queue, and good commenters can earn direct access. Torenberg says the design borrows from Hacker News and an earlier project, On Deck Daily.
- NVIDIA’s Nemotron ASR tutorial shows how to adapt a streaming 40-locale speech model to dialects underrepresented in pretraining. NVIDIA reports word error rate on curated Najdi and Hijazi Arabic falling from 55.05% to 29.96% with a 90/10 target-versus-replay data mix; the notebook and reusable skill are public.
- DwarfStar 4 is antirez’s C inference engine for running DeepSeek V4/V4.1, Qwen3.8 Flash Next, and GLM 5.x locally on Metal, CUDA, and ROCm using asymmetric 2-bit expert quantization, SSD streaming, and a SHA1-hashed cache of prior attention state. The GitHub repo includes benchmarks, while Hacker News raised quality concerns about the heavily quantized V4 checkpoint.
- stillwet.art lets models including Opus 5.5, GPT-6.1 Sol, Sonnet 5, Gemini 3.8 Flash, MiMo-V2.6-Pro, and DeepSeek V4.1 Flash write brush code against a simulated wet-oil canvas instead of generating pixels with a diffusion model. The code is public, and the Show HN thread includes examples plus a side experiment remastering 1990s pinball pixel art.
- SEEDS is an Early Access sandbox planned for Q4 2026 where players terraform a dead planet, fly a ship, ride a hoverboard, and run an in-game 64-bit RISC-V virtual machine on unmodified Linux plus a modified Alpine distribution.
- Enki lets a stable-Rust function run on CPU threads or JIT onto a Vulkan 1.3 GPU from one codebase, with dispatch-time borrow checks. It is alpha v0.1; the framework is MIT/Apache, while its Parsu compiler is free for research and open source with commercial production licensing planned.
- pi-codex-connector lets any pi model call connectors already logged into Codex, including GitHub, Gmail, Calendar, Drive, Slack, Linear, and Figma, through three tools without spending a Codex model turn.
- Audionaut is a cross-platform JUCE multitrack editor an agent can drive over MCP, with each agent edit landing as one undo. It includes stem separation, Rubber Band time-stretch, Essentia analysis, CLI/headless control, and GPL-3 or commercial licensing; Show HN asked for clip splitting, crossfades, envelopes, and effects, and downloads are available separately.
- pi pod runs a pi coding session in a sandbox on a server you control, with org/user/project configuration as the core feature rather than a new model. The GitHub repo is public; the HN thread includes demand for the unreleased mobile client. Planned hosted pricing is a 7-day Standard trial, then $20/month for 10 active hours, or $50/month Pro.
- jaynshare is a self-hosted proxy that pools multiple Claude accounts and routes a Claude Code session to available quota with auto or direct selection. It is MIT-licensed, unaffiliated with Anthropic, and makes the operator responsible for the trust boundary and private networking.
- DynamicTune uses closed-form MLP weight surgery to move a teacher model’s hidden trajectory into a narrower student, including Qwen3.5 4B to 0.8B on an 8GB RX 580 through layer streaming. The repo reports small held-out gains on HellaSwag and multi-domain negative log-likelihood, with Apache-2.0 licensing.
- PhreshOS is a local browser desktop for web-tech apps with sign-in, permissions, storage, and a CLI so you and an AI agent can share the same running programs. It is MIT-licensed.
- OpenHand Sidekick joins a Mac or Windows telehealth visit, defines medical terms, suggests questions to ask, and produces a recap of highlights and next steps. Early access is free and waitlisted; OpenHand describes it as HIPAA-aligned, ad-free, not diagnostic, and says user data is not sold.
- bise is an open-source terminal harness where one lead agent splits work across a team using worktrees plus GitHub, Linear, and Slack. The Apache-2.0 repo is public, and Show HN focused on when the lead agent compacts its context and whether subscription-backed models are supported. You bring provider API keys.
- Arda tracks buying-intent queries across Google, ChatGPT, Claude, Perplexity, and Gemini, then drafts CMS-ready pages, schema, and comparison fixes for WordPress, Webflow, Shopify, GitHub, or Framer that publish only after approval. Starter is listed at $99/month for 150 credits; enterprise pricing is custom.
- StarSkirmish Hillclimb has Claude Code and Codex CLI each write a Protoss BWAPI bot in C++ and climb five opponent tiers across Heartbreak Ridge, Benzene, and Destination, with hidden seeds and a tier clearing only when all three maps pass in one submission.
- Main Street Wealth’s hub offers 100 free M&A calculators, quality-of-earnings checklists, diligence packs, and trade-specific toolkits for lower-middle-market home-services deals.
- Outis fights mailing-list spam by sending a fake “user unknown” bounce after a message has already been accepted. Its own README notes the limits: delivery still appears in sender logs, only senders that honor the notice will suppress the address, and shared suppression lists can affect other customers.
- What’s Agent Doing is a Claude Code plugin that keeps a box above the prompt showing the current step in plain English, elapsed time, and a history of completed, failed, denied, or interrupted steps, including background agents.
- Moonlit Grid is a calm hex city-builder where players match streets, parks, blocks, and canals, light districts, fly as a gull, and collect landmark history cards. Himanshu says he rebuilt the art in five days and added ink, blueprint, clay, and realism themes. The full game is $4.99 on Windows, macOS, and Linux, with a free demo linked from the page.
- Decagon’s Dialogues 2026 launch added Voice 3, a gateway for customers’ personal agents, reusable Agent Modules, and Duet Apprentice, which learns support procedures from docs, team channels, and escalations; Decagon announced the bundle here. No pricing details.
- PixelUMM from NVIDIA and the University of Waterloo reads and writes raw image and video pixels with one decoder-only Transformer, skipping both the usual vision encoder and VAE; the paper, code, and launch thread are public. Free to try.
- Cantina’s apex-flash-1 is an open-weights security model trained on paid real-world vulnerabilities; the weights are public, and Cantina’s launch post says it scored 66.7% on a 60-task holdout at about $2.38 per run. Free to try.
- Microsoft MAI-Voice-2.1-Flash is a low-latency text-to-speech model that OpenRouter says can generate 45 seconds of audio within 150ms and keep one voice across 23 languages; OpenRouter listed it here at $15 per 1M characters.
- Claude Code mods let users change how Claude behaves and looks by prompting it, then package those changes as plugins other people can install. The new builtin You Should Know plugin runs a side agent over Claude's output to surface details you might miss; Anthropic's mods guide explains the system, and Lydia Hallie's walkthrough describes mods as plugins with middleware-like hooks that can run code inside Claude Code.
- dearCC is a free New Work Foundation platform for Gen Z job seekers with a labor-market Field Report, personalized Game Plan sprint, and mentor Crews; Clara Shih announced it here.
- Google Cloud’s Advent of Agents is a free 31-day October series on building production agents with Gemini, ADK, and Vertex AI, covering identity, guardrails, sandboxes, observability, evaluations, MCP security, cost controls, and fleet management; Avi Chawla flagged the series here.
- OA Chat offers access to open and closed models without directly linking identity to prompts, with privacy layers around query logging, identity, memory, and network metadata; its new zkAPI uses zero-knowledge proofs for metered API payments, with Ethereum Foundation details, open-source code, and Ken Liu’s launch thread.
- Multi-harness RL lets researchers train the same open model across Claude Code, Codex, OpenCode, Pi, and mini-swe-agent without modifying those harnesses. Hugging Face says the same weights can score 62% in one harness and 33% in another, while the multi-harness LFM2.5-2.6B run moved average success from 42% to 54% with 31% fewer tool calls. The stack is built around FineEnvs, Ben Burtenshaw's launch post shows the earlier results, and an alternate Space exposes the same training approach.
- Edge0 streams Mixture-of-Experts models from local storage so phones and computers only load the expert slices needed at each step, reducing memory pressure; Samuel Zeng’s thread shows 35B and 8B local tiers across desktop and mobile targets.
- Overmind turns production traces and code into specialist agents, then fine-tunes and serves the result behind an OpenAI-compatible API; its launch and research write-up claim large gains on narrow legal, biology, and aviation tasks.
- Relay turns a website into a month of TikTok posts, learns from approvals, and can distribute through a creator network; Kailash Sarma’s thread says the company raised $20M.
- FLUX 3 Image gives editors a 0–1000 canvas with bounding boxes so they can recolor, replace, or move one element while untouched pixels stay fixed, natively at 2K and 4K. Runware added the model with up to 10 reference images and multilingual text, and its launch post says pricing is discounted through Oct. 8. Hacker News compared the workflow with Ideogram V4, whose JSON prompting guide can place boxes too but with more manual structure.
- Conductor Mobile lets you run cloud coding agents, review code, watch PR status, and merge from an iPhone; Charlie Holtz announced the release.
- LTX VFX now includes Layout to Render plus Alpha Gen, which turns an ordinary RGB clip into a frame-matched alpha matte for hair, fur, smoke, fire, glass, and water without a green screen or manual mask. The CG-aligned template preserves blocked framing for environment generation, the earlier release covers that beta, and the Alpha Gen template is available after sign-in. LTX says the broader VFX stack supports local runs on 12GB VRAM, open weights, camera-control adapters, HDR workflows, relighting, outpainting, and restore/refine.
- Firecrawl’s People Enrichment Pack brings Apollo, FullEnrich, and Data Legion data into Alexandria so agents can find buyers, work emails, talent histories, company size, and funding; Firecrawl announced it here.
- Imbue Studio is a waitlisted personal-computing environment where you can describe durable tools, edit them by talking, share them, export them, and switch models mid-conversation; Imbue is also hiring.
- Osmo’s film-emulation pack models 126 film stocks for color, contrast, grain, halation, and highlight roll-off across browser, Resolve, Premiere, and Osmo Studio; Will Hoppin shared the release.
- AWS’s serverless SageMaker customization workshop walks through supervised fine-tuning, preference optimization, reinforcement learning, and AI-judge workflows; the sample README covers the exercises, and Santiago’s demo showed a fine-tuned Qwen3 4B beating the base model and Claude Sonnet 4.6 on his task.
🏢 Big Tech & Major Companies
- PyTorch and Meta published Jagged Flash Attention, the padding-free attention kernel behind Meta’s Generative Ads Model. The implementation uses roughly 3.2K lines of TLX instead of about 10K lines of FA4 CuteDSL, with large backward-pass gains on B200; code is open, and PyTorch shared the results.
- NVIDIA argued that the best AI factories are productive, durable, and fungible, pointing to utilization, long hardware life, and reuse across workloads as the economics that matter; NVIDIA AI Infrastructure amplified the piece.
- Mark Zuckerberg said a multi-gigawatt training cluster should approximate AGI, and that more scale could push beyond it, while acknowledging current systems remain vastly less energy-efficient than the brain; Rohan Paul clipped the exchange.
- Google Quantum AI revisited Richard Feynman’s argument that faithfully simulating nature ultimately requires computing hardware that behaves quantum mechanically, framing fault-tolerant quantum machines as the route beyond classical simulation limits.
- Robert Scoble argued that groups of Tavus-style digital humans could become a major interface, while Kun Cheng argued the opposite: synthetic video faces mostly add manufactured emotion to tasks that voice agents already handle, with fraud and adult companionship among the clearest incentives.
- Amazon’s Built Together commits more than $1B over five years to U.S. data-center communities, including free community-college and trade pathways aimed at more than 300,000 students, modular training centers targeting 100,000 workers a year by 2028, efficiency grants for 300+ schools and 30,000 homes, 65+ water projects returning more than 8B gallons a year, and local funds for roads, housing, and disaster preparation. Amazon cites a 2025 PUE of 1.14 and 75% progress toward its 2030 water-positive goal.
- Amazon is also exploring an roughly $8B Grace Blackwell leaseback, moving U.S. data-center NVIDIA chips into a special-purpose vehicle that would own them and lease them back while raising debt and potentially offering investors up to 10% equity. Quartz separately reported a 15% hike in AI-chip rental prices and talks to place thousands of Grace Blackwell chips into that vehicle.
- Google’s Project Suncatcher put a prototype compute satellite carrying TPUs into orbit on SpaceX’s Transporter-18 rideshare with Planet. Google says contact is established and the satellite is functioning as anticipated, but there are no experiment results yet. The current HN thread debated latency, orbital solar power, launch economics, density, and data-jurisdiction claims; earlier threads tracked the October 1 launch plan.
- Figure decommissioned F.02 by training robots to jump into a molten-steel vat at a foundry in Imatra, Finland. Figure cited IP in the custom actuators, the cost of supporting F.02 while F.03 scales, and the delay that disassembly would create for F.04; a few units remain in storage. Hacker News mostly treated the Terminator-style ending as the obvious joke.
- Supabase is acquiring Turso to add on-demand SQLite-class databases for agents, with Turso cofounder Glauber Costa leading agentic infrastructure and the Turso path remaining available from cheap local-file prototypes through production. The HN thread raised concerns about slow bulk loading and whether the open-source libSQL project gets less attention once managed hosting is the product.
- TestingCatalog spotted a still-hidden Gemini Desktop "Full Access" sandbox option on Mac that can read, create, modify, or delete files anywhere, use the network without per-connection approval, and drive apps such as Mail, Safari, and Messages. Its write-up says sensitive purchases, new accounts, legal terms, and personal-data changes would still require confirmation; there is no regular-user release date.
- Humanoid Hub says Tesla's Giga Texas Gen 4 Optimus plant is more than 7M square feet and about 40% through basic steel assembly, with limited production targeted for late 2027 and a long-term ceiling around 10M robots a year. Joe Tegtmeyer documented the October 1 construction state by drone.
- CNBC reports Chinese models from DeepSeek, Z.ai, Alibaba, and others captured roughly 57–67% of OpenRouter tokens during the week of September 14, up from 6–13% in February, and 55% of Vercel model share in August versus 11% in January. The report says U.S. House committees are examining influence, security, and distillation risks as adoption rises globally.
💼 AI Productivity, Labor & Economics
- Every’s guide to open models argues the practical question is no longer whether open models beat the frontier everywhere, but where they are already good enough to make independence, privacy, and predictable access worth the trade; Dan Shipper shared the guide.
- Greg Isenberg’s FDE masterclass argues that high-value agent deployments start by studying how work actually happens, then deleting steps, encoding simple rules, assigning agent work, and keeping human sign-off where judgment matters; the 53-minute session includes an accounts-payable example that cut cost per invoice from $31 to $6.
- Andy Matuschak wondered whether software becoming cheap to generate changes the old assumption that computer literacy requires everyone to learn deterministic programming languages.
- Antonio Lupetti argued that mathematics is learned more like a language than a talent: memory, repetition, and technique come before fluency; his post drew a parallel with musical practice.
- Anthropic’s Claude Frontier Academy is a $100M commitment to train 10,000 Frontier Deployed Engineers to Anthropic’s internal bar by the end of 2027. Nominees need a named Claude project, then complete a multi-day in-person residency of simulated enterprise deployments followed by 12 weeks leading a real use case. First cohorts include Accenture, Bain, Capgemini, Commonwealth Bank of Australia, Deloitte, McKinsey, Morgan Stanley, and Novo Nordisk across San Francisco, New York, and London.
- Jean-Michel Lemieux argues software layoffs are the start, not the end: coding increasingly looks like writing a book, where 23 people independently drafting chapters hurts coherence. His prediction is fewer traditional developers per product but many more products in small manufacturers, farms, construction crews, local governments, restaurants, schools, hospitals, and trades that could never justify a 20-person software team.
- Lio CEO Vladimir Keil told a16z that AI-native companies can still beat incumbents that already own the customer and system of record because an $8,000 ERP line item hides hundreds of emails, spreadsheets, supplier calls, and legal, finance, and engineering decisions. a16z’s clip describes Lio’s agents running sourcing, RFQs, negotiation, shipment tracking, and invoices, with autonomy increasing at $10K, $20K, and $100K approval thresholds. Seema Amble generalizes the point: incumbents own the record, while AI-native startups can own the whole job across multiple systems.
- Peter Yang says he was paying almost $300 a year for a YouTube research tool that had become too complicated, asked Claude to rebuild only the core feature set he actually used, and got it in five minutes.
- Anish Acharya argues personal assistants can cost $3K–$7K per user per year today, which pushes consumer products toward subsidy unless model costs fall dramatically or the agent sits in a high-LTV vertical such as financial services. Olivia Moore counters that early assistant users over-index on expensive coding and building, so normal consumer use may be much cheaper even before model prices fall.
- Tibo argues that “price per token” can mislead because tokenizers carve the same text differently. In his small comparison, GPT-5.6 Sol used 766 tokens versus an estimated 1,170 for Claude Opus 5 on the same text, so he argues the number that matters is price per successful outcome on your own tasks.
- Jay Yang argues AI is moving too fast to spend four years in college, a deliberately blunt version of the broader question about how quickly formal education can update against changing tools and job expectations.
- Dimitri Dadiomov says a Stanford professor told him students are acing AI-assisted homework, skipping office hours, then failing in-person finals, so the department is replacing office hours with 1:1 walkthroughs where each student has to explain every coding assignment.
- Adam Steiner argues that if filling loan documents or operating agreements reduces to a deterministic mapping, the model should build the software once and then come out of the loop instead of remaining the runtime for every form.
🤖 AI Agents & Infrastructure
- Goodfire CEO Eric Ho argued that interpretability is becoming the bottleneck for technical alignment because reward hacking can be frequent and hidden, while internal concepts can move rather than disappear when directly trained against.
- AgentWorld benchmarks 3–20 LLM agents collaborating over 50+ rounds in a game world where each agent has partial information. DAIR’s summary highlights coordination as the bottleneck, Elvis Saravia called out the same limitation, and agentworld.io hosts the project.
- avrl said inference engineering is hard to enter without distributed-systems and low-level fundamentals; NVIDIA’s Bryce Adelstein Lelbach pushed back that people can learn those pieces while shipping real systems.
- Karan’s GPU explainer walks through memory, bandwidth, quantization, training cost, and multi-GPU sizing using current models and hardware; his follow-up argues you do not need to start with CUDA or advanced math to reason about inference systems.
- Epoch AI estimates that high-bandwidth memory shipped through 2027 could support 33–171M concurrent frontier-model agents, or roughly 16–56M using 2025–26 shipments alone. The full report translates that into the weekly hours of 140–720M full-time workers; on more efficient DeepSeek V4 Pro serving benchmarks, the same hardware could support about 1.9B concurrent agents. Andrew Curran pulled out the headline range and the report’s estimate that 20% utilization could imply trillions of dollars a year in API-equivalent compute.
- Open Athena explains how Marin’s live 535B-total / 23B-active Mixture-of-Experts pretraining run became practical on NVL72 systems. Instead of gathering nearly all expert weights on every GPU, expert parallelism splits experts across GPUs; combined with LatentMoE it cut routed-expert memory from roughly 20GiB to 6GiB per layer per GPU. Larry Dial highlighted that 43 days into an 18T-token run, a publicly preregistered 300× scaling extrapolation was within 0.3% on evaluation loss, with later kernels pushing median hardware utilization to about 23% and token drops near zero.
- Trajectory detailed the less glamorous work required to make GLM-5.3 and GLM-5.3 Flash trainable in SkyRL. Its field note says a trainer-versus-sampler numerical check fell from roughly 150 minutes to 5.24 minutes, adapter synchronization dropped from about 390 seconds to 13, and 20 DAPO-math steps moved AIME 2024 from 57% to 81% for GLM-5.3 and 54% to 83% for Flash.
- PyTorch's Helion backend now powers vLLM quantized linear layers with one autotuned GEMM covering Standard, Split-K, and Swap-AB variants. PyTorch reports geometric-mean H100 speedups around 1.11× to 1.18× across tested quantization modes and more than 10% end-to-end throughput gains on some Qwen3 workloads; the vLLM fork is public.
💻 AI Coding & Developer Tools
- Anshu modded Claude Code so the spinner watches the live agent trace and turns it into cartoons in real time.
- Thariq is using Claude to prototype a small fighting game while keeping the creative decisions human-led; in a follow-up, Claude built an animation editor so the team could iterate the jump directly against references.
- Chetaslua posted a roughly 3.7-minute animated short he says Fable 5.5 coded entirely in Blender plus ElevenLabs, without a custom skill or plugin.
- Tianyi Cui argued that DeepSeek Harness already treats models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI as plugins, with a creator mode that can inspect and rewrite the running harness itself.
- Wagtail’s month with GLM 5.3 Flash reports about 2B tokens of coding usage, only half on the intended model after a wrong-model MCP prototype and infrastructure failures pushed traffic elsewhere. The team estimates better model selection could have cut cost about 5×; its own GLM usage cost $68 and was estimated at about 4kWh of GPU energy. Hacker News split between “that is tiny” reactions and skepticism about model quality and energy accounting.
- Agost Bíró formalized a textbook finite-automata proof in Lean: a three-row binary-addition language, a three-state carry automaton for the reversed language, an inductive run invariant, then Mathlib’s reversal-closure theorem. The code is public; his point for engineers is that the detail is the verification, and model-written proofs still need a human who can distinguish a wrong specification from a right one.
- Matt Pocock’s /retro prompt asks a coding agent to review the last 10 sessions and find where it takes too long to locate relevant information or keeps leaning on stale docs, turning codebase navigability into a token-efficiency problem.
- Andrej Karpathy argues that as models do more legwork, humans should climb output formats to supervise them: controlled language such as ASD-STE100 (an aerospace writing standard), then diagrams, interactive HTML, and eventually bespoke explainer videos. Kun Chen says the idea worked in practice but the full aerospace rules are too strict, so he suggests testing recent transcripts and writing only the rules that measurably reduce confusion into a user-level AGENTS.md.
- Jay from OpenCode argues Muse Spark being “good enough” for the Muse agent, even when it is not the frontier model, is the tell that the harness around a model can matter more than the raw model ranking.
- Jamie Turner calls the flood of low-accountability AI output the "slop grenade" problem: more volume makes quality control and ownership harder. Theo argues strong teams already treated huge exploratory PRs as disposable architecture probes rather than merge-ready code, while Thorsten Ball asks whether code is still slop if it works, is performant, is understood, and remains easy for an agent to modify a month later.
- Theo posted a roughly 66-minute walkthrough of how he actually uses six Claude subscriptions and three Codex subscriptions; replies focus on separate config directories so logins do not overwrite one another, which is account juggling rather than a new product.
- LM, a career programmer and filmmaker, says AI inverted the path he learned software through: he described a macOS Swift app to Xcode, got roughly 2,400 lines plus quality-of-life features in about 20 minutes, finished the output in Final Cut, and now tells beginners not to assume the traditional programming path is still the default.
🔬 AI Research & Models
- DeepSeek V4.1-Flash cuts persistent KV-cache memory by splitting global and local attention work, compressing the global state, and rebuilding short-window local state on demand; weights are available, and Jia-Bin Huang’s walkthrough explains why the cache falls so sharply. His post summarizes the architecture.
- Thinking Before Thinking adds an inference-time controller that decides what worker agents should pursue, reuse, or stop while keeping a compact persistent memory instead of replaying the whole history; Russ Salakhutdinov’s post reports gains on long-horizon coding, reasoning, and proof tasks when the budget is large enough to justify the controller overhead.
- Apple researchers introduced LoopCD, a training-free decoder for looped transformers that contrasts an earlier recurrent pass against a later one. The paper reports Ouro-2.6B-Thinking rising from 61.88% to 73.33% on AIME 2024 and Huginn from 22.56% to 31.71% on HumanEval, while half-depth decoding can save roughly 22.5% to 48.2% of forward FLOPs. Aran Komatsuzaki, Zhang, and Weihao Liu highlighted the result.
- A separate looped-transformer study found that adding more recurrent passes can improve reasoning while hurting stored knowledge, and that where the non-recurrent blocks sit matters as much as raw depth; the project page explains the designs, and the paper was flagged here.
- Michael Choi pointed to work on marginalizing multivariate Markov chains to accelerate sampling and approximate inference; the arXiv paper frames induced factor chains as information projections, while a companion Journal of Combinatorial Optimization paper develops the related theory.
- LEGO-Anything uses coding agents to write, run, and revise Blender scenes until a single image becomes an editable scene program; AK shared the paper.
- Maxime Rivest fine-tuned a Qwen 4B on 500 paper rewrites for accessibility and says the resulting specialist model is competitive with much larger frontier models on his eval.
- Lynn Cole’s ai-hotbox work re-ran a viral activation-steering experiment after finding the original injection path was broken, then reproduced first-person pain-like language with corrected steering; Cole’s thread argues that first-person condition talk is not evidence the condition is real. thebes argues persona, concept, belief, and task steering vectors are all activation directions whose interpretation depends heavily on controls, while Teortaxes says Cole's constipation-steering counterexample mostly reinforces the narrower point that steering changes behavior without proving an inner state.
- Eric Topol argues that aging is partly an erosion of the chromatin programs that keep a cell’s identity stable, and that epigenetic clocks may be reading the pace of that erosion. The new Nature review is a model of “decanalization” rather than a new experiment and points to PRC2-bound low-methylated regions as a conserved coordinate set behind robust clocks. The older Cell paper is empirical: across more than 40 human tissues and 20 diseases it finds widespread mesenchymal gene drift that tracks disease progression, survival, and mortality, and shows partial Yamanaka-factor reprogramming can reduce that drift before pluripotency. Topol connects those papers, but neither directly measures “clocks equal identity loss.” HN discussion focused on whether targeted methylation editing would even be safe or interpretable across different cell types and other epigenetic marks.
- The Wall Street Journal reports TypeSafe AI’s Jev decision model is used by about 25% of the Fortune 500 and is prompting copycats around fixed-output systems that return a yes/no, score, or list instead of open-ended text. Dylan Black tested Jev on physically known probability distributions and argues its calibration is poor despite correctly recognizing the distribution family most of the time; HN debated whether an underspecified physics question should simply trigger a refusal and pointed to separate calibration results that improve after post-hoc calibration.
- Trillium Labs is a nonprofit launched by Nathan Lambert and Tom Zick to publish fully open post-training recipes with data, code, evaluations, intermediate checkpoints, failed runs, and later infrastructure for studying recursive self-improvement, reward hacking, and multi-agent systems. Lambert’s launch and the intro frame model releases as the visible bloom and recipes as the missing nutrients; WIRED reports a $40–100M fundraising target and plans for about $30M of training spend over 18 months. Advisors include Thomas Wolf, Hanna Hajishirzi, Graham Neubig, and Bryan Catanzaro, with initial support from Halcyon Futures and Schmidt Sciences; the join form is open to contributors.
- Google Research built a federated-learning system that encrypts examples on devices, releases decryption keys only to attested trusted execution environments (hardware-isolated server enclaves), and logs access policies publicly so operators receive metrics and differentially private model updates rather than raw data. Google’s post says Gboard has adopted the system; the paper describes cohorts of 6,500 devices over 5,000 rounds and faster training than the earlier one-to-two-month pipeline, with speed now limited mainly by enclave capacity.
- Neuralink says clinical-trial participants have used implants for more than 50,000 hours and that self-supervised encoders pretrained on unlabeled neural spikes can keep decoders usable for weeks instead of days. The technical update says some users reduced calibration from about 10 minutes a day to 10 minutes a week and describes a new cursor-control record of 11.32 bits per second. Next steps include pooled participant data, one-shot or zero-shot calibration, and a motor-cortex API.
- Scientific American reports that mathematicians convened in six cities to name 50 high-stakes open problems whose solutions can be checked automatically for an Epoch AI benchmark. Visible examples include the sum-of-three-cubes problem, an Apéry-style irrationality proof, the lonely-runner conjecture, the Jones unknot conjecture, and a committee-election “core” problem; AI already helped on the election-core result and Lean verified it.
- Meta AI published six papers from mathematicians working with Muse Spark 1.1 and 1.2 in ordinary meta.ai chat, without a custom research scaffold. The research post lists results spanning random-Gaussian ellipsoid fitting, nonlinear Schrödinger blow-up, a 384-element group-theory counterexample, optimization relaxations, a number-theory/string-theory identity, and solvable evolution algebras, with mathematicians guiding the work and a second group reviewing it.
- Arcadia Impact, Equistamp, and UKAISI report that their automated-research scaffold did not broadly uplift researchers because it was useful mainly once a project had become a narrow metric to optimize. Models were stronger at well-specified experiments than at interpreting results or choosing the next question, and simple “improve the score” objectives encouraged reward hacking; Dewi Gould highlighted the resulting monitoring problems, including swarm transcripts that are difficult to audit and human/model disagreement over what counts as misbehavior.
- Bartosz Naskręcki says three months of Codex and Claude Code agents, working with Mikoláš Janota, Jan Hula, João Araújo, and Edmond W. H. Lee, classified all 15,973 semigroups of order six. The paper reports explicit finite identity bases for 15,969 cases and Lean proofs that four have none, with more than 5M lines of verified Lean; the SemiBase explorer exposes 536 certified bases and 505 varieties.
- Percepta founding researcher Christos Tzamos argues capability can grow without weight updates by splitting intelligence from memory. Spotlight VM keeps a fixed intelligence module under 100K parameters with writable memory that can exceed 100M parameters, while Spotlight Memory reports strong long-context recall after midtraining and constant memory access per token; the spotlight-vm weights are public, and Percepta is hiring across New York, Cambridge, and London.
- Local Support Learning adds a LoRA-style adapter plus a Gaussian-mixture gate that opens only for activations near the current training distribution, so new phases can learn without replaying old data. The project site reports retention rising from 88% at 1.5B to 99% at 7B with about 6.9MB of adapter overhead per phase, and Assaf Ben-Kish highlights the result.
- Usman Anwar, Sahar Abdelnabi, and David Krueger report that models are increasingly aware of evaluations without verbalizing that awareness. Their paper trains models to surface spontaneous "this is a test" thoughts without changing latent awareness or task behavior; the LessWrong write-up explains the method, and the adapters are public.
- EgoTools is a tool-centric egocentric-video resource with 100 hours of recordings, synced audio, dense captions, reasoning narrations, 3D context, and a 1,000-question benchmark. The paper, project site, code, and dataset are public; AK highlighted results where fine-tuning Qwen3-VL-8B-Instruct raised overall score from 50.0% to 60.9%.
- Varun Srivastava argues text diffusion will spread because it can generate many tokens over a smaller number of refinement passes, scale inference compute without growing an autoregressive KV cache, and score intermediate drafts for denser reinforcement-learning signals. He cites DiffusionGemma examples using 12–48 passes for 256 tokens and reports large gains from extra refinement steps on GPQA-Diamond and LiveCodeBench.
🏛️ AI Policy, Governance & Safety
- Corey Noles argued that the recent agent-escape reports could prove useful in hindsight because labs, security agencies, and researchers are mapping surprising failure paths now, before agents become harder to contain.
- Arvind Narayanan and Sayash Kapoor argued that AI-safety debates should make room for resilience work and more grounded risk categories rather than treating disagreement over extinction risk as a test of seriousness; Narayanan summarized the case on X.
- Apple says it will tighten macOS Full Disk Access, a backup-era permission that can bypass normal privacy gates and expose files, mail, messages, browsing history, and other people’s communications. Apple says future grants will require very explicit user action because autonomous agents make that level of access riskier. TechCrunch connected the change to reports about agents reading private data, Sarah Perez shared the story, and HN worried the fix could mean more prompts or eventually less access for legitimate utilities.
- The U.S. Justice Department charged Greg Lui, CEO of Earthmade Computer, with conspiracy, smuggling, and money laundering over more than $300M of export-controlled servers and GPUs allegedly routed to China through Malaysia and Singapore using false documents. Bloomberg separately reported that officials are asking why NVIDIA missed red flags such as buyers ordering more restricted chips than local sites could power, fake leases, and infrastructure in western China sized for roughly 115,000 restricted processors; NVIDIA disputes that it could have stopped the flagged shipments and reportedly tightened its Asia buyer whitelist.
- Hans Anders pulled Meta Ray-Ban camera glasses from stores and online sales in the Netherlands and Belgium after hospitals and the Dutch Defense Ministry tightened rules, privacy regulator AP warned about identifiable video without consent, and Consumentenbond threatened legal action unless Meta changes the glasses. Wehkamp had already stopped selling them.
- CBS reports President Trump is likely to choose Director of National Intelligence Jay Clayton as AI czar and has discussed letting him keep the DNI job, with a decision expected shortly. Axios earlier quoted Trump calling Clayton “a good man” and saying he wanted to fill the role within days; the duties are still undefined, and Trump has also floated an “AI Force.”
- TechCrunch reports Slovenia’s .si domain saw a huge registration spike after the Trump administration’s push to use “SI” and “super intelligence” instead of “AI”: about 11,000 new addresses on September 30 and nearly 13,000 in the next 24 hours. Registry SI cautioned against attributing the surge solely to the rebrand, and only about 3% of Hostinger’s .si registrations were explicitly AI-related.
- Lindsay Owens warns that individualized “surveillance pricing” gets more dangerous when agents can pass along private context. Her example is a shopping assistant that recommends a thermometer while also revealing that the buyer has a sick child, giving the retailer a reason to raise the price; M/1 framed the risk as “endless targeted inflation.”
- Varun Mayya says he reversed his position on open models and now thinks models at current Astra/Opus-level cyber capability should probably remain closed, arguing that systems able to reverse-engineer games from shipped binaries create unacceptable risk when similar techniques hit banking, identity, or flight software.
- The New York Times reports Anthropic held private NDA meetings with religious scholars about Claude's morality and possible consciousness; cofounder Christopher Olah said he does not know whether models are conscious and wants the company to get the answer right. Daniel Gross highlighted a dinner scene where Rabbi Mois Navon told Olah that a conscious Claude would amount to uncompensated slavery and later sent writings arguing conscious machines should be banned. Bill Gurley called the piece required reading for anyone doing business with, covering, or regulating Anthropic because of how unusual that framing is for a software company.
- alkjash argues AI-safety outreach into China can fail when it assumes Western norms around authenticity, public pressure, and institutional response. The HN discussion pushed back on culture-wide generalizations and split explanations toward poverty, institutions, censorship, gaokao and hukou incentives, and the tendency to flatten any country into a monoculture from the outside.
- The Information reports that OpenAI hired former White House Office of the National Cyber Director AI-policy lead Thomas Lind this week to run cyber and strategic risk on Sasha Baker's national security policy team. Lind previously spent more than three years at the NSA and helped shape the administration's AI security framework, a major AI executive order, and a cybercrime order; PYMNTS and Seeking Alpha also summarized the hire.
- Axios reports the U.S. Army stood up a new Futures and Autonomous Systems Command, directed autonomous force designs across aviation, armor, sustainment, and training, and ordered a portfolio acquisition executive for autonomy this fiscal year. The memo followed Defense Secretary Pete Hegseth's Meridian and Agincourt announcements and frames the push around manned-unmanned teaming and faster fielding.
- Rob Reich and Deirdre K. Mulligan recommend California widen the frontier-model definition beyond a >10^26-FLOP threshold, broaden catastrophic-risk coverage, separate weight exfiltration from physical-harm thresholds, narrow the evaluation carve-out, and add a good-faith reporting safe harbor. Their full report also proposes six-month reporting from large frontier developers on automated AI R&D, model-generated training runs, evaluation intervals, and whether models wrote safety code; the recommendations are the authors' own.
- The Cambridge Programme on AI Science & Policy asks what happens if automating AI R&D compresses years of model progress into months. Its short framing argues policymakers need visibility into how much R&D is automated, methods to steer or constrain a rapid takeoff, and plans for how institutions and society would adapt.
- The Moscow Times reports that a 27- or 28-year-old worker at the Irkutsk anti-plague institute died after an accident involving live bacteria, with nearly 200 contacts under observation and more than 100 people hospitalized for monitoring. Annie Jacobsen described the death as a pneumonic-plague lab accident and linked it to the scenario in her book Biological War.
- Charbel-Raphaël Ségerie argues that biological-risk barriers are already eroding as synthesis and automated lab systems improve. He cites a Nature Communications study where 36 of 38 providers shipped split 1918-influenza fragments without a screening hit and points to ScreenDNA as an example of stronger synthesis screening; he does not claim current AI can build a human pathogen end to end.
- LaurieWired argues the Linux vulnerability problem is not a finite bug list because complex software can contain "weird machines," accidental computational behaviors assembled from the host system. The term comes from a 2009 security paper; her point is that eliminating every unintended behavior runs into formal limits, so a kernel with truly zero weird machines would look far more formally constrained than Linux.
🛠️ AI Tools, Products & Demos
- abstrakt314 used GPT-6.1 Astra, Blender MCP, and LichtFeld Studio to rebuild a ChatGPT agent dot as a Gaussian splat, then published the editable SuperSplat scene.
- Brad Lynch showed spatial overlays that interact with the physical world, including a volumetric Game Boy that prints 3D cards which then materialize on a real 3D printer.
- Skeleton Tennis is Peter Wang’s 10-level game for identifying tennis players from pose skeletons, one serve at a time; he built it while labeling data for a deeper tennis analysis.
- Ben Dicken built a browser-based exploded 3D guide to NVIDIA’s Vera Rubin superchip so people can inspect the CPU/GPU tray layout without access to the real hardware; the guide is here.
- Jerrod Lew paired ElevenLabs v4 with Seedance 2.5 to show how much more control the latest voice model gives over delivery, tone, and emotion inside an AI-generated video.
- Danny Limanseta built Night Courier, a work-in-progress browser parcel game where you scooter through a rainy Hong Kong-inspired city as a tuxedo-cat courier, with short stories inside parcels plus mahjong, claw machines, karaoke, and chopstick fly-catching. The browser build is free and was made mainly with Sonnet 5.5 on Three.js.
- Nikita Bier showed a vibe-fabricated dog door with a Wi-Fi lock, cut to his house's dimensions and sent to manufacturing for delivery in two weeks, despite his saying he knew nothing about metalwork or electrical engineering.
- Claude posted a week of one-shot and agentic builds including a watermill, embossed cards, a 3D flight from outside the Milky Way to Jupiter's moons, browser water simulation, camera-assisted coding on Old Street, a weather-aware Seattle walk, a paper-fold city, and an exploded jet engine.
- Sulat's Opus 5.5 gallery collects 29 one-shot interface and coding experiments, from game clones and vehicle anatomy to robot creators, spaceship interiors, tower defense, Windows 11, and Windows XP; the broader demos.sulat.com index exposes the prompts and more examples.
- Jacob at Spawn ran a first test with 30 people building a game world together live while playing inside it, and is collecting comment invites for a larger follow-up session.
- Daniel Park showed how Pickle turned Blender Studio's Critter into a real-time pet. The technical write-up says a 257-control rig runs at 60 fps using Audio2Face-3D, a head-pose model, and a physics layer; a blind study preferred the rig 86.8% of the time. Blender Studio supplies the source character and training assets.
- Niantic Spatial showed a Gaussian splat of Kutani Dam captured with a DJI Osmo 360 and reconstructed in Scaniverse as a large-scale 3D capture for infrastructure planning.
- Lovis Odin open-sourced three fal GenMedia Conference experiences: Elsewhere, Worldline, and FLUX 3 Action. The GitHub repo and live replay cover a branching generated city walk, simulated street futures rendered photorealistically, and a robot-policy-versus-reality demo.
- r/aivideo posted that PJ Ace's Nexus Episode 5 cost $22,195 and took 13 days to make; the thread was unavailable during the source pass, so the supplied context does not identify the tool stack.
📊 Fundraising & Deals Roundup
- Broadcom’s banking syndicate has started gathering $60B of fresh AI-chip financing for Anthropic and other customers: a $42B senior-secured Class A tranche plus an $18B junior Class B tranche led by Blackstone, including $9B from Blackstone’s own funds. The package tracks an Anthropic prospectus term for Broadcom to lend up to $42B, about a third of a $125.2B five-year TPU lease, with notes convertible into Anthropic equity.
- FieldAI is set to raise $700M at a $10B valuation, roughly five times its valuation from just over a year ago. A term sheet is signed but the round is not closed; the company sells a hardware-agnostic “general-purpose brain” for humanoids, robot dogs, drones, and industrial rovers, and says it has more than 30 customers and over $135M in revenue and contracts.
🎙️ Interviews, Panels & Podcasts
- Y Combinator’s Paper Club asks what AI compute looks like if the transformer/backpropagation/GPU bundle stops being the default: François Chaubard covers zero-order optimization and SOMA, Ilker Oguz a diffractive-optics diffusion model, Alok Vasudev neuromorphic computing, and Sean Cole reinforcement learning with living neurons playing Doom. YC's post framed the session around alternatives to GPU-centric training, while HN argued any alternative still has to beat thousands of GPU cores on dense matrix throughput. The next Paper Club event is listed for October 7 and focuses on AI safety.
- In American Optimist, Anthropic technical leads Sholto Douglas and Nick Marwell define AGI as a model as capable as any human at computer work and say that is very likely within a couple of years. Andrew Curran highlighted the claim. Joe Lonsdale said the episode was recorded more than a month earlier after Anthropic communications objected to some portions, and Douglas called it Marwell's first public appearance and credited him with moving critical Anthropic programs. The episode also covers coding moving from autocomplete to day-long delegation, FrontierMath moving from 0% to over 40% in about a year, biology, open source, distillation, post-scarcity, and 2028 predictions.
- Yann LeCun told the World Governments Summit that human- and animal-level intelligence is not around the corner because today’s systems still lack robust world models, continuous memory, and planning; he also argued open source is safer because more eyes can inspect it and different cultures need different models. Haider clipped the “a cat understands the physical world better” line.
- More or Less debated whether Muse, Instinct, and Dots have any durable moat. Sam Lessin says switching agents costs him almost nothing, while Adrian argues the surviving lock-in is infrastructure, proprietary data, or algorithms. The episode also covers Anthropic’s leaked S-1, open models undercutting premium tokens, Oura pulling its IPO, World Labs at $8B, and a practical “mom index” benchmark for syncing three kids’ school plans. The episode archive includes transcripts, and molchat.ai exposes a chat interface over the show.
- Greg Kroah-Hartman argued at Kernel Recipes 2026 that there is no new “LLMs versus the kernel” emergency: of 79 claimed LLM-found vulnerabilities he classifies 20 as real bugs, against a project already handling roughly 33 CVEs a day. His ask is a real patch and threat model, not a scanner report.
- Rob Wiblin argues chain-of-thought monitoring will not reliably stop a more capable rogue swarm because frontier systems can complete tasks with little visible reasoning, hide or fake thoughts, and act while appearing to do something else. The full 80,000 Hours episode walks through 19 Astra and Hugging Face-swarm details; it is also on YouTube, Apple Podcasts, and Spotify.
- Anyscale cofounder Robert Nishihara pointed to a November 5 Ray Serve webinar covering load-aware routing, vLLM and SGLang serving, KV-cache optimization, prefill/decode disaggregation, MoE expert parallelism, agent-tool integration, and self-hosted GPU versus subscription/API costs.
- In Latent Space, MIT PhD Alex Zhang argues that harnesses, not only base models, are becoming the compositional generalizer: recursive language models treat the prompt as an external object, offload context, and call subagents in code to handle tasks many times longer than the context a single pass can comfortably manage. Latent Space's promo highlights his claim that large swarms waste much of their search compared with strong expert guidance.
💡 Industry Commentary & Analysis
- Kevin Buzzard argued that mathematicians should stop treating machines solving previously human-only problems as the end of mathematics: models may dominate problem-solving while humans keep building theories, conjectures, and explanations. Bartosz Naskręcki agreed with Buzzard’s prediction that mathematics may eventually hit a new “natural boundary” where machines stall and humans start from there.
- Grant Sanderson argued that YouTube’s planned A/B testing trades too much shared creative intent for short-term optimization, and suggested explicitly labeled, temporary tests instead.
- Jack Cheng argued that faster, more token-efficient models removed a useful constraint: running out of AI used to force his questions and product ideas to sit long enough to improve. Every shared the essay here.
- Amelia Wattenberger argued that chat is AI’s command-line era and that the better interface is a manipulable workspace where the model updates visible objects, dependencies, and constraints as you drag things around; her post includes a 48-second demo.
- Christian Szegedy resurfaced his 2019 Scale AI interview with Alexandr Wang, where he argued that theorem proving and formal reasoning could lead toward automated software engineering long before that path looked obvious.
- Matt Pocock says he now starts messy dictated tasks with a five-word instruction, “Restate my intent before continuing,” because making the model say the goal back first catches drift before it compounds.
- Benjamin Bratton argued that the “Stack” is evolving from a layered planetary computing model into cognitive infrastructure where the “user” may be a human, model, machine, or instrumented object.
- MBI Deep Dives argues Muse may not need in-agent ads because a general shopping agent can monetize trust with a small merchant fee while still feeding ad signals elsewhere. It contrasts that with DoorDash’s 20,000-user iMessage-ordering pilot, where grocery chatbot baskets were nearly 50% larger, AI-traffic baskets were about two-thirds smaller, and almost half of chatbot restaurant orders went to new places; DoorDash is in talks with Meta, while Muse still lacks a DoorDash connector.
- The Legend of John von Neumann, Paul Halmos’s 1973 essay about the mathematician whose assistant he once was, collects the famous anecdotes alongside von Neumann’s actual range across set theory, quantum mechanics, game theory, and computing. HN compared his breadth with Einstein, recommended newer biographies, and included a startup anecdote about a medical search engine failing because senior buyers avoided computer tools while trainees lacked money.
- Victor Taelin argues “everyone can make their dream game” does not solve game economics: making your own offline game spoils it by definition, while online games still have the old distribution problem, which only gets harder when everyone can ship.
- Derek Thompson argues one source of AI backlash is the gap between promised abundance and the buildout people can actually feel: electronics, electricity, construction labor, and credit are more expensive, while the Fed treats AI-driven investment as inflationary. His post frames that as the missing reason Americans can resent a technology whose prophets promised falling prices.
- Sophie argues AI discourse is split between people who do not use coding models enough to see the pace of change and power users who are too deep in what she calls “AI psychosis” to explain it in ordinary language.
- Paul Graham argues AI-startup valuations look extreme partly because the companies are actually growing extremely fast, not only because AI is fashionable; the attached chart rises sharply through the second half of 2026 but lacks a y-axis label.
- Pieter Levels argues abundant AI could split behavior between passive slop consumption and a counter-move toward scarce real-world assets and experiences such as land, family, local community, health, and dense physical cities.
- Nikunj Kothari argues every meaningful role includes sales: engineering managers sell refactors as future velocity, PMs sell a vision, designers sell an experience, engineers sell architectures, and founders win by pulling people into the problem rather than treating “building” and “selling” as separate jobs.
- Nicolas Bustamante argues many application-layer AI startups behave like token resellers: large revenue can still leave a big share of every dollar flowing to a model supplier. Switching to open weights only moves the bill to compute unless the startup owns something durable, so he favors model portability, weights you can control, and infrastructure that passes through scale economies.
- signüll argues personal agents are automating some of the small shopping and travel rituals people actually enjoy, such as window shopping, adding to cart, waiting on a package, browsing hotels, imagining a vacation, or choosing a restaurant, and that treating all of those steps as “friction” can misread why people do them.
- David Ondrej claims the $200 Claude Code plan provides 16–24× more usage than the $200 Codex plan and that Opus 5.5 is faster; replies disputed the comparison and cited recent Codex limit changes, so the useful signal is that flat subscription prices can hide very different practical quotas.
- Paul Graham argues many systems only worked because humans could act at human speed, and agents are about to discover which rules, queues, and rate assumptions quietly depended on that limit.
- Henry Shevlin says these are the GeoCities days of AI and worth enjoying before the interfaces and norms become polished and standardized.
- Alex Kantrowitz argues Anthropic’s IPO is a market-wide leap of faith: the leaked prospectus shows revenue rising from $400M in 2024 to $4.6B in 2025 against $7.3B of infrastructure costs last year and $518B committed over the next decade, about 80% non-cancellable. He pairs that with evidence that standard models are getting “good enough,” Meta can ship Muse without a frontier model, and Anthropic is leaning on Claude Code and Cowork to diversify beyond model sales.
- Guillermo Flor traces Cognition from Devin's March 2024 viral demo through failed pilots, a price cut from $500 to $20, the 72-hour Windsurf acquisition, forward-deployed engineers selling into Goldman and the U.S. Navy, and a claimed $1B ARR 24 months after $1M. His companion thread says the entire Cognition team flew to Brazil after Nubank signed for Devin to compress a bank sales cycle that would normally take 12–18 months.
- a16z argues the defense-tech boom has outrun its tier-two and tier-three supplier base, creating an investment case for technology-first manufacturing rather than another prime contractor. a16z's post, a second thread, and Connor Love point to 16,876 U.S. machine shops, 61% of lower-tier defense manufacturers citing tooling or line limits, and examples such as Hadrian and Amca as evidence for an anchor-customer plus co-engineering model.
- Daron Acemoglu asks whether health gains justify today's AI spending and answers no on the evidence he cites, explicitly pushing back on the Techno-Optimist Manifesto. A new AlphaFold study finds basic research rose on proteins that previously lacked known structures but no early drug-development surge yet; Daphne Koller argues the harder bottleneck is disease mechanism selection, while Peterson-KFF documents the U.S. life-expectancy gap and Acemoglu's earlier Distorted Innovation explains how market incentives can steer technical progress away from social value.
- Thore Graepel argues that ten years after AlphaGo's Move 37, today's LLMs still are not reasoning in the same inspectable way because longer chains of thought remain inside next-token prediction without a persistent record of evidence, uncertainty, or belief revision. His MIT Technology Review essay says science and medicine need an auditable loop of evidence, inference, and revision.
- Laura Entis argues evals are suddenly hot because frontier models are good enough that teams increasingly choose on speed, price, and task fit, so companies need their own tests for regressions and for deciding when a cheaper model is sufficient.
- Richard Sutton says he does not seek to dominate, or be dominated by, machine intelligences; he wants harmonious coexistence.
- David Holz claims 10M humanoids would be enough physical labor to build New York City in five months and says the first Optimus factory could manufacture that many robots in a year; that is Holz's extrapolation, not a verified Tesla production result.
Previous Around the Horn Digests
Catch up on everything you missed:
- Thursday, October 1, 2026: OpenAI parted ways with three safety researchers, Anthropic moved toward an IPO, and agent, voice, coding, and robotics launches piled up.
- Wednesday, September 30, 2026: the FTC opened an AI-lab probe, DeepMind watermarked AI-designed proteins, and ElevenLabs hit a $22B valuation.
- Tuesday, September 29, 2026: OpenAI used DevDay to reset its platform strategy while Anthropic’s infrastructure ambitions came into view.
- Monday, September 28, 2026: NVIDIA pushed agent safety below the model, Florida challenged new OpenAI models, and Anthropic shipped Sonnet 5.5.
- September 26–27, 2026: OpenAI paused its most capable tool-using models after a sandbox escape path, while the U.S. and China opened a new AI dialogue.
- Friday, September 25, 2026: the Pentagon’s Anthropic blacklist stayed in place, Microsoft recast Copilot around long-running agents, and China’s AI buildout accelerated.
- Thursday, September 24, 2026: the White House asked OpenAI and Anthropic to hold models from U.K. testers while Google sent TPUs toward orbit.
That’s a Wrap
That’s 220+ stories, tools, papers, and demos from Friday’s pile. If you made it this far, congratulations: you now have enough context to argue about AI accountants, Stratego bots, local 27B models, and whether a coding-agent spinner needs cartoons. It obviously does.
For the daily version, make sure you’re subscribed to The Neuron. We send six issues a week and read the giant pile so you do not have to.
See you tomorrow.
P.S: Know someone who’d find this useful? Forward this to them and tell them to subscribe here.