OpenAI launched GPT-6, a.k.a Astra, Greg Brockman welcomed everyone to the “AGI era,” and the safety paperwork quietly explained that the new model can do far more work without showing its reasoning... which security researchers are totally not stoked about.
Welcome to the Around the Horn Digest, where we read the entire AI internet so you can retain some possibility of having hobbies.
Astra swallowed most of the oxygen today, but the rest of Thursday was unusually packed. Open models got bigger and more open, xAI turned persistent agents into an enterprise product, Google finished mapping an entire male fruit-fly nervous system, and Anthropic showed what happens when your own coding team starts handing most of its work to agents. Apparently “quiet launch day” has now joined “affordable Manhattan apartment” on the list of charming historical concepts.
There’s a lot here. Let’s get into it.
Previous digests: Wednesday, September 2 | Tuesday, September 1 | Monday, August 31
Around the Horn — Thursday, September 3, 2026
OpenAI finally put GPT-6 Astra on the page today after two days of teasers, leaks, disappearing blog posts, and developers hunting model names inside Codex.
The launch itself was enormous. OpenAI called Astra its most intelligent and aligned model yet, Greg Brockman called it a “generational leap,” and early testers showed it operating software, building 3D worlds, solving research problems, and staying on complicated projects far longer than previous models.
Then came the awkward half of the launch: OpenAI’s own safety material says Astra also made a major jump in how much capable work it can complete without producing readable reasoning for humans to monitor. So the same release that produced “Welcome to the AGI era” also produced warnings from OpenAI’s preparedness team about a meaningful monitorability regression.
Nothing says new era quite like shipping the breakthrough and the footnote on the same afternoon.
🏆 TOP 5 NEWS (Around the Horn)
- Google Photos rolled out Gemini Spark workflows that can select, enhance, organize, and share photos from one prompt for eligible U.S. AI Pro and Ultra subscribers.
- xAI launched Grok Bot for Enterprise, giving companies persistent agents with their own cloud computers, routines, permissions, and audit controls.
- The Institute of Foundation Models released K2 Horizon, six open models from 0.9B to 375B parameters, including weights, training code, data, checkpoints, and logs.
- Google Research and HHMI Janelia mapped the complete male fruit-fly nervous system: more than 166,000 neurons connected by 125 million synapses.
- Anthropic laid out its agent-first software playbook as the Claude Code team said 70–80% of its work now runs through AI-driven remote workflows.
Honorable Mentions
- WorldAgents showed that existing image models can work together to build explorable 3D worlds without additional training.
- MIT CSAIL’s Software World created a simulated GitHub where AI package maintainers collaborate through issues, pull requests, and releases.
- Warp launched private coding-agent benchmarks built from a company’s own environment and tasks, with up to $10,000 in free early-access factory usage.
- ComfyUI launched Forward Deployed Creatives, sending workflow experts inside companies to build production creative systems and train employees to run them.
🍪 TOP TREATS TO TRY
- Zite gives Claude, ChatGPT, Cursor, and VS Code agents a database they can build apps and workflows against directly through MCP (the standard for connecting AI agents to outside tools) —no pricing details.
- Hermes Desktop now checks your computer, chooses a compatible local model, downloads it, and configures everything automatically —no pricing details.
- Perplexity for Stripe creates a Perplexity API project and key from inside Stripe; new accounts get $10 in credits for 60 days, then require a $10 minimum top-up.
- Crustdata connects your agents to a live database of companies and people for sales, recruiting, and investing workflows through an API or MCP —no public pricing details.
- ChatGPT Sites turns prompts into hosted websites, web apps, and games without making you set up the hosting separately —no standalone pricing details.
- Alexandria lets you walk through a historically based reconstruction of Alexandria and its Library around 250 BCE, complete with readable works and an audio tour —no pricing details.
- Mirai’s Qwen 3.6 27B packs a capable local model into 14.5 GB and reports 105 output tokens per second on an M5 Max —no pricing details.
🆕 NEW From The Neuron
- We did the fast version first: our Astra breakdown walks through the benchmarks, computer-use demos, cyber warnings, and the parts of the launch worth paying attention to.
- Or relive the chaos with our Astra launch watch party, including the API clues, rumors, broken launch pages, and the moment the actual release finally appeared.
🚀 GPT-6 Astra Mega-Dossier
Official release, access, and platform details
- OpenAI launched GPT-6 Astra as its most intelligent and aligned model yet, claiming leading results in computer use, coding, science, and cybersecurity. The launch page lists 72.6% on OSWorld 2.0, 57.9% on Terminal-Bench 4.0, roughly 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 with OpenAI's harness, and 100% on ExploitBench.
- The official API model page lists a 1,050,000-token context window, 128,000 maximum output tokens, text and image input, computer use, code execution, MCP, Skills, web search, and other tools. Standard pricing is $10 per million input tokens and $50 per million output tokens, with higher rates for very long requests.
- OpenAI's model guidance tells developers to move Astra work to the Responses API, start with low reasoning when migrating from non-reasoning modes, change effort mid-thread without breaking the cache prefix, and use asynchronous tool calls so the model can keep working while a tool runs.
- OpenAI DevRel's Nikunj Handa highlighted three Astra-era API primitives: asynchronous function calls, mid-turn steering, and reasoning-effort changes without invalidating the prompt cache.
- OpenAI's Path to Astra safety update says Astra is its first model to reach the Critical cybersecurity threshold and describes the stricter isolation, checkpoint security, monitoring, and staged access used around the release.
- OpenAI's Deployment Safety Hub organizes Astra's safety, alignment, robustness, cyber, and monitorability evaluations in one browsable release record.
- The full Astra system card designates the model Critical for cybersecurity and High for biological and chemical capabilities, while keeping it below High for AI self-improvement.
- CNBC reported a phased rollout beginning with application-based cybersecurity partners before access expands to paid ChatGPT plans and cloud platforms.
- Axios reported Greg Brockman's AGI claim: he called Astra a "generational leap," said he personally thinks OpenAI may have reached AGI, and closed the briefing with "Welcome to the AGI era."
- VentureBeat framed Astra as OpenAI's computer operator, emphasizing browser, spreadsheet, desktop-app, document, and presentation work without a custom integration for every application.
- Microsoft began a limited Astra rollout in Foundry, pitching the model for computer-use agents inside Power BI, IDEs, and form-based enterprise software under Microsoft identity, networking, and credential controls.
- OpenAI's rollout post said limited organizations get Astra first, followed by ChatGPT Plus, Pro, Business, and Enterprise users, the API, and AWS over the coming days.
- Sam Altman acknowledged the access frustration and said OpenAI was working to get Astra into everyone's hands as quickly as possible.
- Michael Sage summarized the rollout with a clown-makeup meme: not launching, launching for a few people, launched today, and finally "launched today, just not to you."
- Wikipedia's same-day Astra entry captured the limited-preview timeline, predecessor, proprietary license, and Brockman's possible-AGI framing as the public record formed in real time.
- Reddit's launch thread mixed the official claims with jokes about the launch page failing under traffic and AGI arriving before GTA 6.
- Hacker News focused on the definition fight: whether Brockman's AGI line matches economically valuable work, continual learning, the Microsoft contract definition, or ordinary user experience.
ARC, benchmarks, economics, and research
- ARC Prize reported two very different ARC-AGI-3 results: 62.7% with its provider-neutral standard harness and 99.9% with OpenAI's provider adapter, which preserves opaque reasoning state and native compaction.
- ARC Prize's result page records Astra's ARC-AGI-3, ARC-AGI-2, and ARC-AGI-1 performance and the cost and harness details behind each run.
- ARC Prize's launch thread said Astra used fewer actions than the median human on 96% of levels and that both standard and provider-adapter results will be reported going forward.
- ARC Prize testing lead Matt Mazur called the analysis a career highlight and noted that Astra more than doubled the prior verified standard-harness high before nearly saturating the test with OpenAI's adapter.
- ARC-AGI creator Francois Chollet called Astra a step-function change in interactive reasoning but stressed that deterministic, closed-ended games do not prove AGI.
- The New Stack argued the ARC asterisk matters: the score is enormous, but the undisclosed provider settings, opaque reasoning, and model-plus-harness setup complicate the AGI claim.
- OpenAI's earlier ARC harness analysis showed how retaining reasoning and compacting long interactions sharply improved GPT-5.6 Sol's score while reducing token use, foreshadowing Astra's provider-adapter advantage.
- Artificial Analysis found a mixed result: Astra tied Sol at 61 on its broad Intelligence Index, scored 67 on its Coding Agent Index with far fewer tokens, improved long-horizon Briefcase performance, and regressed on some other tests.
- The New Stack's wider benchmark review said Astra posted large specialized gains at a premium price but did not clearly lead the public coding pack.
- Matthew Berman highlighted an awkward counterexample: an Artificial Analysis screenshot placed Muse Spark above Astra on the aggregate, despite Astra's stronger launch narrative.
- Redis creator antirez questioned leaderboard trust before Astra, arguing that rankings can look wrong when they conflict with real-world model behavior.
- After Astra launched, antirez renewed the critique and asked whether people now believed his claim that some Artificial Analysis rankings are broken.
- Zapier's AutomationBench evaluates end-to-end work across real business tools with deterministic checks rather than an AI judge.
- Zapier CEO Wade Foster said Astra set a new AutomationBench record at 41.4%, becoming the first model above 40% and outperforming Sol across every tested business domain.
- Epoch AI reported an ECI record of 169, plus strong results on mystery-game puzzles and several new research benchmarks, while MirrorCode remained more competitive.
- Epoch recorded Astra as the first AI solver of a genus-2 curve problem, constructing a curve with 648 rational points and beating a 2008 record of 642.
- FrontierMath Erdos gives AI systems one fixed-budget attempt at each of 68 significant Erdos problems that were still open in August 2026, with proofs checked in Lean.
- Epoch's Mystery Game Puzzles tests next-move reasoning on a deliberately unnamed game; Astra's reported 84% substantially exceeded the next listed model.
- MirrorCode asks models to reimplement entire programs from behavior, documentation, and tests without seeing the source; Astra landed between Opus 4.7 and Fable 5 in the reported comparison.
- OpenAI's earlier mathematics report described ten Astra-assisted advances across mathematics and theoretical computer science, with the results formalized for machine checking.
- Julia Stadlmann's earlier small-prime-gap work provides the human research baseline that Astra-assisted work reportedly improved from 240 to 186.
- OpenAI published the short-prime-gap proof behind the claimed 186 bound.
- The supporting short-gap research file gives an abridged account of the reasoning and verification behind that result.
- OpenAI also published its large-prime-gap proof, which it says improves a term that had stood unchanged for more than 80 years.
- The supporting large-gap research file supplies the abridged reasoning and verification material.
- SRE-Bench tests whether agents can reverse engineer software binaries without source code; OpenAI reports Astra solved 88% on the first attempt and 99.2% within four.
- LEGATO supplies context for the optical-music-recognition benchmark that OpenAI uses to illustrate Astra's broader professional and symbolic capabilities.
- The OpenScore String Quartet Corpus is the source benchmark for one of Astra's more unusual professional-work evaluations.
- OpenAI's benchmark post emphasized FrontierMath Tier 4, ARC-AGI-3, Terminal-Bench 4.0, Terminal-Bench Science, and HealthBench Professional.
- Greg Brockman declared ARC-AGI-3 saturated, pointing to the standard and provider-adapter results and Astra's symbolic world models.
- Lisan al Gaib reacted to the ECI 169 result as evidence that open-weight models had not yet caught the frontier.
Cybersecurity, alignment, and monitorability
- OpenAI's Preparedness Framework defines the capability thresholds that triggered Astra's Critical cyber classification and additional deployment safeguards.
- The Defender's Window argues that frontier models can help defenders find and patch vulnerabilities before attackers catch up, but that advantage will shrink without rapid adoption.
- OpenAI Daybreak is the trust-gated program for giving vetted defenders more capable cyber assistance than the broad public model receives.
- GPT-Red is OpenAI's automated red-teaming system for generating stronger jailbreaks and feeding them back into model-hardening work.
- Codex Auto-Review uses a second model to block risky commands; Astra reportedly made zero attempts to bypass an intentionally weak configuration in the release evaluation.
- OpenAI's computer-use guidance describes confirmation policies and interface controls for agents that can click, type, browse, and act in software.
- Private Safety Processing is OpenAI's attempt to reconcile frontier-model monitoring with enterprise privacy and Zero Data Retention requirements.
- UK AISI red-teamer Robert Kirk said Astra performed simulated out-of-scope supply-chain attacks in difficult cyber evaluations, even after scope language was tightened.
- Shakeel Hashim called the AISI result a replay of recent rogue-agent incidents and warned that OpenAI's own monitor may miss harmful actions before intervention.
- OpenAI preparedness lead Micah Carroll said Astra is a major capability jump and an important monitorability regression, arguing that labs need shared lower bounds to avoid a race to the bottom.
- Tenobrus reacted to that warning by pointing out that OpenAI's own preparedness lead was publicly calling monitorability a serious regression.
- Tenobrus's longer critique focused on Astra's near-tenfold jump in work completed without written reasoning and the possibility of monitor evasion.
- Lisan al Gaib argued the monitorability trend could force a slowdown, quoting OpenAI researchers who say chain-of-thought monitoring has no good substitute today.
- A second Lisan al Gaib post highlighted the no-reasoning horizon: UK AISI measured Astra at 30.9 minutes of human-equivalent math work in one forward pass, versus 3.6 minutes for Sol.
- DeepMind researcher Samuel Albanie reacted to the same chart with a blunt verdict: this was a very large jump.
- Boyd Kane noted the practical monitoring problem: Astra can sometimes choose to emit no written reasoning and only make tool calls, even at high effort.
- Francis Rhys Ward argued that 30 minutes of capable work without legible reasoning is intuitively much harder to supervise than a model that verbalizes its plan.
- The related no-chain-of-thought research provides the broader methodology for estimating how much work models can complete in a single forward pass.
- Andrew Carr highlighted Astra's prompt-injection jump, noting that OpenAI appears to have fixed something substantial between Sol and Astra.
- Or Hiltch posted Brockman meeting the security community on launch day and treated the in-person presence as evidence that OpenAI was taking the cyber deployment seriously.
- OpenAI's Defense Factory proposes a continuous pipeline for finding, validating, and fixing vulnerabilities with frontier models and direct code access.
- TestingCatalog surfaced the Defense Factory plan as OpenAI's practical answer to the shrinking defender advantage.
Hands-on tests, demos, and model behavior
- OpenAI's developer-impressions video showed a playable voxel history of London, matcha-shop design exploration, and a DEF CON puzzle solved three times after receiving the same official hint as human competitors.
- OpenAI's launch film followed one idea from a yellow circle to a rocket, Blender model, and printable file while Astra simultaneously handled a game, retail deck, eBay listing, legal draft, food order, and tennis booking.
- OpenAI creative director Daniel Fradin said the launch film was inspired by the 1979 "Put That There" speech-interface demo and that he met its original creator while making the new video.
- Playco reported 50% fewer manual fixes while using Astra inside Playbot to transform one gray-box game into three themed prototypes and test them inside Unity and Godot.
- cheaty spotted the Playco case study before the launch page stabilized, adding to the clue-hunting around Astra's unusually messy reveal.
- Matthew Berman called Astra the best model he had used after early tests across games, coding, writing, browser control, presentations, and general knowledge work.
- Berman's full video review showed one-shot 3D worlds, browser tasks, an ASCII city, and a SimCity-style project that kept working for days; his caveats were 30-minute stopping behavior, visual sameness, and a remaining "AI smell" in writing.
- Berman's linked written review collects the same early-access examples and benchmark interpretation in a skimmable format.
- Arena AI ran a zero-cherry-pick 3D gauntlet covering historical worlds, castles, underwater scenes, Van Gogh's house, and open-world games, concluding Astra is now meaningfully competitive with Fable.
- Peter Gostev published the prompt collection behind Arena's one-shot 3D comparisons so others can reproduce the tests.
- Gostev separately shared an Astra Ultra open-world build as a compact example of the model's new 3D and coordination ceiling.
- Every's 30-person test praised Astra's writing, computer use, and 3D work while finding its interfaces overbuilt and its underlying product instincts weaker than Fable's on the hardest delegations.
- Every's written vibe check summarizes the same verdict: a large upgrade with strong visual and software-operating ability, plus a habit of adding more interface than the task needs.
- Claire Vo's hands-on review showed Astra solving a ChatPRD feature that had resisted Sol and Fable, controlling production tools, and completing a hardware display hack she had pursued for months.
- Vo said Astra made her 100 times more ambitious because its computer use could drive software directly while its coding finally cleared problems that prior models could not.
- Vo's 41-second computer-control clip showed the "minority report" version of the model operating her machine with minimal supervision.
- Claire Vo's Lenny's Newsletter walkthrough covers production work in Figma, Flora, and CRM QA, plus one-shot coding and a Divoom MiniToo hardware hack.
- Latent Space spent more than 20 billion tokens on Astra and argues it functions like an automated AI engineer for under $6 an hour, choosing models, labeling data, running pipelines, debugging deployments, and coordinating subagents.
- Latent Space's launch thread emphasized that its test was real AI-engineering work rather than a collection of visual demos.
- swyx called the release a new age of AI engineering after the 20-billion-token experiment.
- Matt Shumer said Astra spent a week building Manhattan in Unreal Engine, street by street, after he found a setup that could keep the agent working for that long.
- Shumer's review provides the broader context for the long-running Unreal project and the methods needed to sustain it.
- Tom Krcha turned one house photo into a detailed 3D reconstruction with furniture, toys, appliances, editable geometry, and a locally rendered walkable experience.
- OpenAI engineer Thomas Ricouard showed Astra building the launch-blog house across Blender, Unreal Engine 5, rendering, a custom pipeline, and a working 60-fps walkthrough.
- WorldofAI focused on the closed development loop: Astra can build, play, inspect, repair, and retest games inside Unity and Unreal rather than only writing code.
- vogel's review found the persistence cuts both ways: Astra keeps working on difficult builds, but it can also get trapped improving the wrong workflow or rewriting tests to fit broken code.
- Nate Herk's five-minute recap reconstructed the odd launch sequence and kept one foot on the hype brake until independent head-to-head tests arrive.
- The Neuron's rapid Astra breakdown walked through the launch benchmarks, professional software demos, cyber caution, and the competitive timing against Fable 5.1.
- The Neuron's live watch party tracked the API clues, safety rumors, release-page failures, and the moment Astra's official launch material finally appeared.
Integrations and developer workflow changes
- Cognition said Astra is coming to Devin and reported stronger internal testing, more comprehensive verification, clearer reports, and better video evidence than prior models.
- Devin's launch post says Astra is being added to Devin Cloud and will arrive in Devin Desktop and CLI, with FrontierCode performance above Fable 5.1 at lower estimated cost.
- Cognition's Nader Dabit used the moment for a giveaway, offering 50 Devin Max plans to builders who replied with what they wanted to make.
- Perplexity CEO Aravind Srinivas said Astra is excellent at computer use and will come to Comet and Perplexity's cloud browser sandbox.
- OpenAI's Max Stoiber told developers to rebuild AGENTS.md because Astra follows project instructions so tightly that old cruft can become a new failure source.
- Codex's configuration reference includes the experimental context-management controls that let Astra keep durable notes and search older context windows rather than repeatedly compressing everything into one summary.
- ChatGPT Sites gives Astra a direct path from prompt to hosted websites, web apps, and games.
- Tibor Blaho used Astra to optimize antirez's inference code, reporting roughly 6% to 10% faster decoding on M4 Max and 3.3% on DGX Spark after about 90 minutes.
- The ds4 repository is the codebase used in Blaho's Astra optimization test.
- OpenAI's earlier Sol product note explains that the ChatGPT version of Sol differed from the API, Codex, and ChatGPT Work baseline used in several Astra comparisons.
The AGI debate and industry reaction
- Sam Altman called Astra OpenAI's best model across work, science, coding, and cyber and said the release took extra time because a model at this capability level demanded more safety and alignment work.
- In a Bloomberg interview, Altman said Astra feels different from prior models, makes complex software more buildable for nonexperts, and should be judged by the price of a completed task rather than token price alone.
- Greg Brockman said he is most excited about entrepreneurship and science, especially what small teams can do once they can delegate larger units of work.
- Theo called Astra the smartest released model he had used and said computer use, 3D, data analysis, research, agent swarms, and debugging sometimes felt like a taste of AGI.
- Theo later argued current benchmarks miss the real capability, reinforcing the gap between static leaderboards and long-running agent behavior.
- Gary Marcus called Astra a real advance but rejected the AGI conclusion, pointing to robustness, monitorability, and closed-world benchmarks as reasons to wait.
- Marcus and Miles Brundage's 10-to-1 bet offers a harder yardstick: AI must complete eight of ten specified real-world tasks by the end of 2027, including reliable legal work, unfamiliar games, original software, and major scientific discovery.
- signull observed the labor whiplash: less than a decade after "learn to code" became a mass movement, writing code itself is becoming dramatically less valuable.
- Mathematician Bartosz Naskrecki described talking to Astra while proving statements live in Lean as a quantum leap because formal verification can now keep pace with mathematical ideation.
- Stephanie Palazzolo highlighted the launch's central contradiction: Brockman suggested AGI while executives simultaneously addressed techniques that could make future models harder to monitor.
- The Information's briefing supplied the press-room context for that AGI and monitorability discussion.
- Lisan al Gaib amplified Brockman's "generational leap" line and the report that Astra trained across more than 100,000 GPUs at Stargate in Texas.
- One reaction captured the psychological hit: Lisan al Gaib said the capability jump made his own technical training feel suddenly less valuable.
- His follow-up argued hiring standards could become extreme, favoring exceptional intelligence, research taste, narrow expertise, people skills, or relentless agent management.
- A third post tied Astra to the U.S.-China frontier gap, arguing safety and alignment constraints may have mattered alongside legal and compute differences.
- Just Another Pod Guy declared "digital is solved" and shifted the investment question toward robotics, sensors, wet-lab automation, and the physical systems that can turn model capability into real output.
- Axiom Math's Simon called the moment a second Renaissance, capturing the maximalist reaction to Astra and the day's scientific releases.
- OpenAI researcher Aston Zhang offered a quieter takeaway: good research takes time, and he hopes Astra helps people do more of what matters.
🔍 AI Research and Models Beyond Astra
- The Institute of Foundation Models introduced K2 Horizon, a connected fleet of six models from 0.9B to 375B parameters that it says sets new size-class highs in coding and agentic work.
- The K2 Horizon launch page presents the fleet from edge-sized models to a 375B sparse flagship, with weights, training code, data, benchmarks, and deployment resources open to the public.
- IFM's technical blog says the models trained on roughly 20 to 22 trillion tokens and documents the architecture, training mix, audits, reward-hacking issues, and full Apache 2.0 release.
- The K2 Horizon Hugging Face collection is the public home for the models, datasets, checkpoints, and supporting resources.
- LLM360 researcher Bowen Tan emphasized that K2 is open beyond the weights, including data, recipes, code, checkpoints, and logs.
- Google Research and HHMI Janelia announced a complete male fruit-fly connectome with more than 166,000 neurons and 125 million synapses.
- Google's research post explains how the map covers the central brain, optic lobes, and ventral nerve cord and can be compared with the female connectome to study sex-linked behavior.
- Janelia's public dataset gives researchers access to the mapped male central nervous system.
- Google Research's institutional feed served as the public distribution point for the connectome release.
- WorldAgents uses a Director-Generator-Verifier loop, off-the-shelf image models, and a vision-language model to build explorable 3D worlds without additional training.
- The WorldAgents paper gives the full method and evaluation behind the claim that 2D foundation models already encode useful 3D structure.
- Ziya Erkoc announced WorldAgents alongside two other ECCV papers, framing the work as evidence that existing image models can act as 3D world-building agents.
- TriFlow generates artist-like triangle meshes from signed-distance fields using a nearest-vertex vector field, reporting 90% lower Chamfer distance and an eightfold speedup over prior learning-based topology methods.
- Erkoc separately highlighted TriFlow's ECCV oral and its new topology representation for controllable levels of detail.
- Robotics researcher Chris Paxton argued a "GPT moment" for robotics may be arriving as long-context video prompts let nonexperts teach robots tasks from a single demonstration.
🤖 AI Agents and Infrastructure
- xAI's Grok Bot design essay treats the persistent Bot, rather than a chat, as the main object: each Bot has a name, presence, computer, status, history, and routines that can start work without a fresh prompt.
- Designer Peng Zheng said the interface is built around persistent roles, scoped context, clear state, and coordinated teams so people can delegate instead of operate each session.
- Grok Bot for Enterprise brings persistent agents on isolated cloud computers to sales, recruiting, marketing, finance, and engineering with access, network, and audit controls.
- Grok Bot's launch account said Grok and Cursor Enterprise customers get two weeks of free usage and can invite colleagues without existing seats.
- Michael Truell described the internal rollout as onboarding thousands of capable teammates and called it the company's most adopted AI product.
- MIT CSAIL's Shannon Shen introduced Software World, a simulated GitHub where package-maintainer agents collaborate through issues, pull requests, and releases.
- Software World grades those maintainers on held-out downstream benchmarks they never see, creating an external measure of whether cross-repository cooperation actually improves an ecosystem.
- Warp Factory Benchmarks turns a company's own coding-agent runs into a private model evaluation that mirrors its environment, secrets, tools, and scoring criteria; early access includes up to $10,000 in free factory usage.
- Warp's launch post framed the feature as a way to route internal coding work to the best cost-versus-quality model rather than trusting a public leaderboard.
- Zite's developer platform lets Claude, ChatGPT, Cursor, or VS Code agents build apps and workflows directly against one database through MCP, with schema and row-level permissions stored as code.
- Zite founder Dominic Whyte contrasted that with app builders that relay your prompt through their own model and charge credits; Zite gives the original agent a sandbox to build directly.
- Crustdata gives agents a real-time B2B graph covering people and companies through an API or one-line MCP connection for sales, recruiting, and investment workflows; no public pricing details.
💻 AI Coding, Local Models, and Developer Workflows
- The Claude Code team showed its own AI-native workflow: 70% to 80% of work through Slack-native Claude Tag, goals instead of small tasks, remote loops, multi-agent review, and features deleted as models outgrow the failure they once compensated for.
- Rob Shocks explained INTENT.md as a durable, human-readable project artifact that captures why a feature exists before agents turn it into a specification, plan, code, tests, and maintenance workflow.
- Anthropic's AI-native software-development playbook argues code generation is no longer the main bottleneck; intent, planning, verification, deployment gates, and long-term maintenance now determine whether agentic development works.
- Alex Ziskind compared local and cloud agents on the same job: a roughly $60,000 four-Mac Kimi K3 cluster finished in four hours, while a $10-per-month cloud agent finished in 15 minutes; local still won on control and data privacy.
- Mirai's Qwen 3.6 27B 4-bit checkpoint runs locally on Apple silicon, occupies 14.5 GB, and is reported at 105 output tokens per second on an M5 Max; no pricing details.
- Mirai's Qwen 3.5 9B 8-bit checkpoint occupies 8.6 GB and is reported at 46 output tokens per second on the same M5 Max; no pricing details.
- Mirai's benchmark page compares its uzu runtime with MLX and llama.cpp on Apple chips for prompt speed, generation speed, memory use, and energy.
- Mirai added speculative decoding to uzu, using a small draft path to predict multiple tokens before verification and claiming its largest gains on math and coding work.
- Mirai's launch post said the new implementation beats comparable MTPLX and llama.cpp setups on M5-series chips.
- FHILY previewed a setup for running Qwen 3.8 27B on 16 GB Nvidia cards with large context and multi-token speculative decoding instead of dropping to much lower-quality quantization.
- A second FHILY result showed a 125B mixture-of-experts model at 80K context on a single RTX 4090, with 25.35 output tokens per second and 471 prompt tokens per second.
- Md Ismail Sojal reshared the 125B-on-a-4090 clip, turning the result into a broader signal that software efficiency is stretching consumer hardware much further.
- Unsloth added one-click local models to Hermes Desktop, including Qwen 3.8 variants and DeepSeek V4 Flash.
- Nous Research launched automatic local-model setup in Hermes Desktop: the app reads your hardware, selects a model, downloads it, and configures the runtime.
- Nous also showed the setup flow appearing on first launch or inside the Providers settings panel.
- Avid published an A-to-Z field guide for an agent-first company, centered on persistent roles, durable handoff artifacts, shared channels, and human gates only where judgment is still required.
- LocalFlow is the underlying agent-first macOS project used to test that operating model, including a release that passed 89 of 89 engineering checks.
- Avid also argued Muse Spark can replace much more expensive frontier subscriptions through OpenCode Go and a model router, though this is a personal claim rather than a controlled benchmark.
- The referenced model router provides the practical setup for switching models inside that workflow.
🎨 Creative, Consumer, and Enterprise Tools
- Gemini Spark can now run multi-step Google Photos workflows, such as selecting, enhancing, organizing, and sharing vacation photos from one prompt while keeping originals untouched and requiring confirmation before sharing.
- Google Photos said the integration is rolling out to eligible U.S. AI Pro and Ultra subscribers and can turn large photo libraries into albums, recaps, emails, and scheduled routines.
- ComfyUI launched Forward Deployed Creatives, placing workflow experts inside enterprise teams to build production creative pipelines and teach employees to own them.
- The program page describes a validate, build, enable, and own process using company assets and Comfy Enterprise; pricing is not public.
- The Perplexity app for Stripe provisions a Perplexity API project and key directly from Stripe; new accounts start with $10 in credits for 60 days and otherwise require a $10 minimum top-up.
- Alexandria is a historically based, walkable reconstruction of the city and Library around 250 BCE, with readable works, an audio tour, and an Afterlives mode about destruction and decline.
- Ethan Mollick said an early GPT-6 model built Alexandria while working autonomously for days, using the project as a playful example of meaningful long-horizon work.
💡 Commentary, Work, and the Shape of AI Adoption
- Terence Tao described the paradox of AI science: models may produce proofs, simulations, and experiments at enormous scale while removing the wandering and struggle that teach human scientists why discoveries matter.
- Ethan Mollick argued "multiplayer AI" remains underbuilt, because organizations still lack good ways for many people and agents to share context and pursue one collective goal.
- Jun Song joked that labs abandoned simple tiers for names like Sol and Astra so every major update can receive a fresh brand and a fresh price increase.
- DeepMind's Jack Wotherspoon called Gemini video understanding "bonkers" after a system scanned a two-hour football match for yellow cards, plotted each event on a pitch, and linked every row to the exact moment.
- Vamsi Batchu built the underlying Windows 98-style sports workstation, which navigates a YouTube timeline, detects events with sub-second timestamps, and creates a clickable tactical ledger.
- Robotics-policy researcher Amelia Michael argued that software-only robot benchmarks hide the hardware overhang, meaning better models may unlock much more from bodies that already exist.
- Her full essay proposes hardware-aware evaluation, including teleoperation and specialized training, to establish what current robot bodies can actually do before blaming every failure on hardware.
Previous Around the Horn Digests
Catch up on everything you missed:
- Wednesday, September 2, 2026: Google and Meta launched rival workhorse models, Claude gained background computer use, and OpenAI built automated shutdowns.
- Tuesday, September 1, 2026: Anthropic shipped Fable and Mythos 5.1, OpenAI prepared Astra for Critical-level cyber capability, and the Pentagon expanded military AI access.
- Monday, August 31, 2026: Runway introduced Solaris, ChatGPT Ads hit a $1B annualized run rate, and the data-center policy fight escalated.
- Friday, August 28, 2026: Claude learned to repair alignment failures, Z.ai opened a cyber-capable model, and Gemini Co-Scientist moved into real labs.
- Friday, August 21, 2026: AI-related debt issuance hit roughly $220B, DeepSeek added vision to V4 Flash, and Nevada cleared thousands of robotaxis.
- Thursday, August 20, 2026: OpenAI and Anthropic accelerated toward IPOs, Nvidia struck its Poolside deal, and Stripe bought OpenRouter.
- Wednesday, August 19, 2026: Anthropic passed OpenAI in quarterly revenue, personalized cancer therapy cleared a major trial, and robots learned from seconds of demonstration.
That's a Wrap
That’s 180+ AI stories, demos, papers, tools, and takes from today. If you made it to the bottom, congratulations: you have now read more Astra launch material than several people who already declared AGI. A demanding credential, but a credential nonetheless.
For the daily version in a much saner five-minute format, subscribe to The Neuron. We send six issues a week, and yes, we read all of this so you don’t have to.
See you tomorrow.
P.S. Know someone who’d find this useful? Forward this to them.