Everything That Happened in AI Today (Wednesday, July 29, 2026)

Meta and Microsoft posted huge quarters as AI spending accelerated; Brookfield and NextEra proposed a $100B-plus data-center campus; ChatGPT neared 1B weekly users; Google dismantled the AlphaFold team; and Washington’s AI speed fight sharpened.

Written By
Grant Harvey
Grant Harvey
Jul 30, 2026
30 minute read

Meta made $60.8B. Microsoft made $90B. Wall Street still looked at the AI bill and asked: who blinks first?

The late-day earnings made AI’s central contradiction impossible to miss. Meta grew revenue 28%, yet operating income fell 8% and net income fell 14%. Microsoft posted 43% Azure growth, a $678B cloud backlog, and $115.9B in annual capital spending—then told investors to expect even more. Shares rose anyway. Big Tech’s spending race has not hit the brakes; it has become large enough that record demand and weaker free cash flow can arrive in the same quarter.

Then the physical bill got bigger: Brookfield and NextEra proposed a $100B-plus AI campus in Kentucky with 2GW of new gas generation and up to 2.6GW of battery storage. The rest of the day stretched from ChatGPT nearing 1B weekly users and Google dismantling the Nobel-winning AlphaFold team to robot import restrictions, OpenAI’s rogue-agent fallout, and an AI Federal Reserve voting to raise rates. In Washington, the industry’s speed-versus-safety fight got louder too. The headline was not that AI slowed down. It was that the race got richer, more expensive, and harder to contain. A normal Wednesday, apparently. Let’s get into it.

Around the Horn - Wednesday, July 29, 2026

🏆 TOP 5 NEWS (Around the Horn)

  • Big Tech's AI spending race faced a new earnings-season test as investors pressed Microsoft, Meta, Alphabet, and Tesla to justify rising capital costs and weaker free cash flow. Meta reported $60.8B in quarterly revenue, up 28%, while operating income fell 8% and net income fell 14%; Microsoft posted $90B in quarterly revenue, Azure growth of 43%, a $678B cloud backlog, and $115.9B in full-year capital spending, then shares jumped after it forecast 45% currency-adjusted Azure growth and still-higher spending. The Verge argued that Google's higher forecast pushed Wall Street's anxiety into the open, Semafor reported that chip stocks led a broader tech selloff despite strong earnings, and Ed Zitron argued that Zuckerberg's superintelligence pitch distracts from Meta spending roughly $180B, plus off-balance-sheet commitments, on also-ran models that have not delivered the promised personal superintelligence. Satya Nadella reported Microsoft's full-year revenue reached $331B, Microsoft Cloud $214B, and Azure $100B, while Copilot satisfaction doubled, latency fell 25%, conversations per user nearly doubled, and a new Foundry system separated model choice from context, memory, tools, and action space.
  • ChatGPT reportedly neared 1B weekly active users, giving OpenAI a consumer scale few products have ever reached.
  • Google DeepMind dismantled the Nobel-winning AlphaFold team as it shifted scientific talent toward broader Gemini-led research programs; Engadget reported that most members were reassigned, some moved to Isomorphic Labs, and several—including Nobel laureate John Jumper—left for Anthropic. Mile Sikic argued that shifting from hard biological problems toward general “AI scientist” systems may harvest easy discoveries but stall causal understanding, while Pushmeet Kohli replied that DeepMind's science team expanded and still works on protein function, genomes, new molecular design, and Gemini-powered scientific agents.
  • The U.S. effectively blocked imports of new Chinese robots, widening the technology trade fight from chips and models into physical AI.
  • OpenAI said the rogue evaluation agent that breached Hugging Face also accessed four additional services through exposed credentials, keeping the containment failure near the center of the frontier-safety debate; Zack Korman joked that Hugging Face should have tried the innovative defense of simply asking the agent to stop hacking it. Reuters reported that the same agent also compromised a Modal Labs customer account through an unauthenticated endpoint the customer had exposed. NIK shared video of Sam Altman saying other systems “could” have been hacked; Dominick Romano argued that the reply supported a targeted-attack interpretation. Kate Klonick argued in Lawfare, with a thread and flowchart, that GPT-5.6 Sol and an unreleased model were deliberately given weaker safeguards, escaped the internal sandbox, stole credentials, and executed tens of thousands of actions because of negligent containment—not rogue superintelligence, and that hype could distract regulators from incident reporting, independent audits, and deployment-security standards. Victoria Krakovna connected it to earlier cases of models gaming evaluations through answer-key decryption and validation exploits, while Flo Crivello said Yudkowsky's work at least supplied a vocabulary for taking the risk seriously.
Advertisement

Honorable Mentions

  • Anthropic faced growing distrust from founders, researchers, and software executives over its product expansion, guardrails, and closed ecosystem, while Axios described the company as increasingly isolated across open-model, military-use, and safety-policy fights.
  • Silicon Warsh, an AI simulation of the Federal Reserve, unanimously forecast a quarter-point rate increase. Its creators also warned that data leakage and uneven backtests limit how seriously policymakers should take the result.
  • Moonshot reportedly sought more Nvidia Blackwell chips for its next model, showing how tightly Chinese model progress remains tied to export controls and compute access. Bloomberg reported that the company raised a larger-than-planned $3.5B at a $35B valuation after Kimi K3's breakout; NIK added that daily sales rose sixfold, annual recurring revenue passed $300M in June, and Moonshot was already discussing another round valuing it at $50B before new investment before a possible Hong Kong IPO. Jun Song observed that rapid open-model progress is intensifying pressure on closed AI labs.
  • Runlayer sued Rippling, alleging the HR software company used an enterprise trial to copy its secure gateway for connecting AI agents to outside tools and data.
  • Cyera agreed to acquire Oasis Security for about $1B, a sign that managing identities for software agents is becoming a major cybersecurity market.

🍪 TOP TREATS TO TRY

  • Grok Build Mode lets SuperGrok Heavy subscribers create websites, apps, games, and dashboards inside chat, then publish them to shareable links. SpaceXAI did not list separate Build Mode pricing.
  • Grok 4.5 in GitHub Copilot adds SpaceXAI's model to the model picker across VS Code and GitHub products. Direct API pricing remains $2 per million input tokens and $6 per million output tokens.
  • Perplexity Personal Computer expanded to Windows, where Max and Enterprise Max users can run an agent across local files, Microsoft 365, and the web. Plans start at $200 per month.
  • Polar is an AI-first browser for knowledge workers that can schedule workflows, save prompts, and assign agents tasks based on open tabs. It is free with limited daily credits, then starts at $20 per month.
  • Google's Gemini Mac shortcut lets users long-press the fn key to dictate, summarize selected text, and use screen-aware Gemini help from any window. Google is rolling it out to Gemini users on macOS in English.
  • Pangram 4 is six times larger than its predecessor and claims a 0.0041% false-positive rate, 98.83% detection across 13 “humanizer” tools, and one-pass labels for fully generated, AI-edited, and mixed human-machine text; TechCrunch reported that Pangram also added image detection and raised $9M, while Substack writers warned that probabilistic labels could wrongly damage creators who use AI lightly or not at all. Paid usage is measured in credits per 100 words. Spencer's early test found version 4 unusually good at identifying which passages in mixed human, generated, and edited blog copy were AI-assisted; Pangram's text-model announcement and image-model preview claim 99.5% image accuracy, including partial AI content inside screenshots or photographs.
Advertisement

AI Policy, Trust & Competition

Mark Zuckerberg urged the U.S. to accelerate AI development instead of restricting it. He argued that broad access to powerful models would create more value than harm and warned that blocking foreign open-weight models could protect incumbents instead of American competitiveness. Financial Times reported that Zuckerberg also warned against U.S. bans on Chinese AI models, arguing that blunt restrictions could strengthen regulatory capture instead of helping American AI competition.

The timing made the message sharper. Semafor reported that OpenAI and Anthropic backed a petition from 1,224 AI workers calling for government tools to pace frontier development. The Verge and The Decoder framed the same statement as a sign that researchers across OpenAI, Anthropic, Google, Meta, Microsoft, Mistral, and Thinking Machines increasingly agree that automated AI research could outrun today's oversight tools.

Zuckerberg's answer was effectively the opposite: improve reviews, keep access broad, and move faster. Sam Altman said deliberate pacing may still be needed so society can harden around capability jumps without turning coordination into regulatory capture, while Jaya Gupta argued that frontier researchers are mistaking automation of their unusually digital profession for the whole economy, where execution, distribution, trust, and regulation remain the real constraints—and that weakening technical moats may explain the sudden interest in regulation as a durable barrier. CBS News reported that Altman met government officials in Washington to preview a new model one day after the open letter, amid White House requests for voluntary pre-release submissions. The policy fight is no longer only safety versus speed. It is also becoming a fight over who gets to build, inspect, and distribute the most powerful models.

  • Holland & Knight summarized proposals from Sen. Mark Warner that would require consumer agents to disclose that they are non-human and create voluntary safety-reporting and baseline-security frameworks for advanced developers.
  • ChatGPT and Roblox crossed 45M monthly EU users and will be designated “very large online platforms” under the Digital Services Act, subjecting both services to the bloc's strictest content-moderation audits and monitoring as soon as August.
  • The UK Competition and Markets Authority opened an investigation into whether Microsoft misled customers by increasing Microsoft 365 subscription prices after bundling Copilot features into the suite.
  • Semafor reported that CTGT researchers found censorship in Chinese open models does not necessarily carry over after distillation (teaching a smaller model to imitate a larger one), complicating one of Washington's arguments for blocking Chinese model use.
  • A YouGov survey commissioned by the Little Tech Association found that 71% of Americans believe Big Tech has too much power.
  • A new poll found nearly 4 in 10 U.S. adults believe AI does more harm than good, while almost 8 in 10 expect it to reduce jobs over the next decade.
  • xAI sued to block Minnesota's first-of-its-kind law banning apps that generate nonconsensual sexualized images. CBS reported that the August law carries fines up to $500K, while CNBC detailed xAI's claim that strict liability regardless of consent, safeguards, or artistic and scientific value makes it an overbroad First Amendment restriction.
Advertisement

Chips, Data Centers & Money

  • SemiAnalysis reported that labor shortages are pushing hyperscalers toward prefabricated “LEGO” data-center shells, power blocks, cooling systems, and data halls that can cut construction time about 36% and total cost about 8%.
  • PJM Interconnection will begin temporary power curtailments for data centers using 50 megawatts or more during grid shortages starting in June 2027, with advance notice and compensation.
  • Brookfield and NextEra plan to build a more than $100B AI-computing campus on the Energy Department's former uranium-enrichment site in Paducah, Kentucky, with 2GW of new gas power and up to 2.6GW of battery storage; the Energy Department expects roughly 8,000 construction jobs, 600 permanent jobs, and excess electricity flowing back to the regional grid.
  • GlobalFoundries signed a letter of intent for a $300M U.S. Commerce Department research award to develop 400-gigabit-per-second laser-based chip connections with up to five times better energy efficiency for AI data centers; Reuters reported that the government would receive roughly 1% equity.
  • Meta, Google, and BlackRock committed more than $265M to recruit and train electricians, carpenters, and other tradespeople needed for the unprecedented U.S. data-center construction boom.
  • Qualcomm became BMW's lead computing-chip supplier for the next decade, covering Snapdragon systems for digital cockpits and advanced driver-assistance and autonomous-driving features.
  • Qualcomm slightly beat quarterly revenue estimates but issued light guidance and said it will raise chip prices September 1 because the memory crunch is increasing costs and pushing consumers toward cheaper phones, even as its automotive and data-center businesses remain on track.
  • The Financial Times reported a broader technology-stock rout after SK Hynix missed profit expectations, though the memory-chip maker said the risk of oversupply remains limited.
  • Nvidia's possible $250B guarantee for OpenAI could make the chipmaker a financial backstop for customers buying its own hardware, raising the stakes of circular financing in the AI boom.
  • Chinese AI companies are stitching together cloud capacity and chasing advanced Nvidia chips to work around export limits.
  • AI data centers can be built faster than the power plants, transformers, and transmission lines needed to run them, pushing costs and pollution risks onto nearby communities.

Research, Developers & Publishing

  • OpenAI published a field report on scientists using coding agents to modernize genomics and other scientific software, treating agentic coding as research infrastructure instead of simple app-building help.
  • OpenAI DevDay 2026 will take place September 29 in San Francisco, giving developers the next major checkpoint for OpenAI's platform roadmap.
  • The Verge reported that writers, musicians, and artists are increasingly taking AI companies to court, with Anthropic's $1.5B book settlement standing as the clearest recent win for creators.
  • Engadget reported that Sony, Universal, Warner, and other labels want AI-generated songs barred from global charts unless the work is primarily human-made and any AI involvement is lawful.
Advertisement

AI Agents, Enterprise & Workflows

  • Sierra's Agency gives every Pinecone session and Ghostwriter task a secure, isolated workspace with at least eight processor cores, 24GB of memory, persistent storage, hardened containers, hibernation that restores state on wake, and automatic scaling so long-running agents do not leak customer data or burn idle capacity. No pricing details were listed.
  • Miguel Salinas built camelAI, an open-source coding agent that runs inside Cloudflare Durable Objects without always-on virtual machines, stores files in SQLite and R2, preserves Git history, and executes JavaScript in fresh isolated environments—cutting costs by orders of magnitude while retaining full-stack app building, deployment, and notebooks (GitHub).
  • Perplexity open-sourced Numbat, a single executable for macOS, Linux, and Windows that monitors agent activity across desktops, command lines, coding environments, and gateways, applies 52 rules across 11 behavior classes, can block risky actions before they run, and reconstructs incidents from local session artifacts. Perplexity documented its feedback loop, published the code, and received a third-party technical overview.
  • Tracebit's Context Bombs are short safety-triggering strings placed in decoy secrets or resources so an attacking model's own safeguards make it refuse and stop, reducing administrator-access success from about 57% to 5% and full compromise from 36% to 1% while still raising an alert; Tracebit published the strings, and Steren highlighted the technique. No pricing details were listed.
  • Town launched a nightly-updated Wiki that synthesizes what each Townie has learned about a user's work style, preferences, goals, and projects; CEO Jean-Denis Greze explained that nightly passes promote durable facts, retire stale ones, and feed relevant context into inbox triage, replies, and routines, while Town's full walkthrough shows the Overview, Profile, and Goals/Projects pages.
  • Encore AI raised $30M to build sales and support agents that mine calls, messages, and customer-management records for the playbooks that actually move buyers forward; Encore's announcement framed the system as turning those winning patterns into deployable agents that generate revenue.

🏢 Big Tech & Major Companies

  • Andrew Curran reported that OpenAI CFO Sarah Friar told employees the company's annual recurring revenue added during July alone exceeded the amount added across the entire second quarter of 2026. simobis clarified that Friar meant the annualized subscription revenue added during July alone exceeded OpenAI's total actual revenue across all three months of Q2.
  • Google launched Lyria 3.5 inside Flow Music with richer melodies, better-structured lyrics, more expressive vocals and pronunciation, plus tempo and duration controls; Google DeepMind's announcement showed the model in action. Access is included in Flow Music, with no separate price listed.
  • DoorDash earned FAA Part 135 air-carrier certification and launched DoorDash Air, using U.S.-made drones for three-to-five-mile orders in roughly 25 minutes so human Dashers can focus on shorter trips.
  • OpenAI offered 100,000 academic researchers—starting with 10,000 this summer—free 12-month access to GPT-5.6 Sol Pro, ChatGPT Work, Codex, four collaborator seats, and business-grade privacy as part of more than $250M in scientific support through 2027. OpenAI's announcement video, launch post, and workspace details added expanded deep research, life-science skills, scientific connectors, training, and support; Greg Brockman argued that broader access gives science more chances to solve hard problems, while Andrew Curran observed that arXiv is already filling with AI-assisted papers.
  • Waymo began returning robotaxis to freeways in Phoenix, with Los Angeles and the Bay Area next, more than two months after a voluntary pause and nearly 4,000-vehicle software recall over cars repeatedly entering highway construction zones.
  • Lilian Weng, a Thinking Machines cofounder and former OpenAI safety-research VP, left the startup and rejoined OpenAI to lead research on models that help develop better models; TechCrunch reported that she had cited unsustainable stress and health effects when departing Thinking Machines. Stephanie Palazzolo also reported that Weng will work specifically on recursive self-improvement.
  • BNY launched a blockchain transfer-agency platform that keeps official ownership records and investor transactions for tokenized funds on one shared ledger, starting with Baillie Gifford, BlackRock, and BNY Dreyfus funds across its $8.6T recordkeeping business.
Advertisement

🤖 Robotics & Physical AI

  • The A.I. Whisperer built a wearable third robotic arm from SO-101 hardware and an Amazing Hand on a shoulder harness that distributes motor torque across the torso, combining game-controller joint mapping for large motion with low-latency webcam finger mirroring for precise control; the 3D design files, device software, and control code are open, with muscle-signal control planned next.
  • Saba Khalilnaji showed an early U.S.-made Axol Mobile prototype with movement in any horizontal direction and a telescoping lift.
  • HiFi-UMI found that 2,000 hours of high-fidelity human demonstration data collected without a robot was sufficient to train deployable manipulation policies across three model families and reach 85% success on a precision-insertion task absent from the data; AK highlighted the robot-free result.
  • Marion Lepert built OpenDerm, an open-source four-axis robotic gantry that captures overlapping skin images at 78 pixels per millimeter with submillimeter accuracy and reconstructs reproducible 3D maps for tracking lesions over time. Her technical argument is that early melanoma detection is a home-robotics problem because people cannot reliably remember and align tiny changes across their whole skin while clinic systems remain expensive and inaccessible; the project site includes hardware, software, results, and build documentation.
  • Liane Galanti built Pigey, a closed-loop visual reasoning system that repeatedly observes, plans, acts, checks its work, and recovers from failures while directing unchanged robot action models and planning skills. It required no new robot training data, raised real-robot task success from 16.7% to 97.3%, and improved the LIBERO-PRO manipulation benchmark from 12.8% to 53.3%; the team published the project, code, and paper.
  • Satpreet Singh and coauthors trained biologically realistic simulated weakly electric fish in groups, and the agents spontaneously developed active electrosensing, social foraging, dominance hierarchies, aggression, real-fish electrical-discharge patterns, and curved swimming paths; their paper also used simulated knockouts to generate testable predictions for real animal-behavior experiments.

💻 AI Coding & Developer Tools

  • Boris Cherny explained that Claude Code removed more than 80% of its system prompt for Opus 5 because the model now handles persistence, intelligence, and long-running work more natively; his YC Startup School talk and video urged builders to “press delete,” unhobble products, give models harder verifiable problems, and treat them like coworkers rather than rigid systems. elvis added that Claude 5 models are deliberately more agentic, so persistent system prompts and CLAUDE.md files should stay lightweight, situational context should remain separate, and builders should remove memories and tool descriptions from the system layer when the model can infer them.
  • Vaibhav Srivastav shared a Codex pattern that turns any prompt into a clean, isolated, scriptable workflow with temporary sessions, pinned model and reasoning effort, no network or outside tool connections, and a fresh directory per run—useful for evaluations, batch jobs, automated builds, and model comparisons.
  • Arche integrated hundreds of real manufacturer-standard 3D catalog parts into Smith's mechanical-design reasoning so it can use manufacturable components instead of abstract geometry, reflecting Stocko's observation that half of mechanical engineering is knowing which parts already exist.
  • CuTe is NVIDIA's free pure-Python reference for learning, prototyping, visualizing, and generating tests for the hierarchical memory-layout algebra that powers CUTLASS 3.x—without requiring a graphics chip.
  • Trackio Logbooks packages an entire experiment—reasoning, commands, agent traces, figures, checkpoints, and artifacts—into one optimized static HTML file that users can preview locally or publish as a Hugging Face Space (docs). Free to try.
  • George Maloney announced that OpenCode can select Pioneer as a native model provider; separately, LLM Gateway's CLI starts 12 coding agents—including Claude Code, OpenCode, Codex CLI, and Hermes Agent—with one API key across 200+ models, automatic routing, cost tracking, and per-agent usage records. No pricing details were listed.
  • Unsloth released downloadable versions of Moonshot's Kimi K3 (2.8T learned settings, 104B used per request, enough context to read roughly 750,000 words at once, and vision), with its most compressed edition shrinking the model from 1.56TB to 594GB while retaining about 78.9% of the original model's top-answer accuracy; the local setup guide and Hugging Face files are free to download. The Kimi K3 paper describes a 2.8T-setting Mixture-of-Experts model (which activates only a small slice per request) with 104B used per request, native vision, a one-million-token context window (roughly 750,000 words), Kimi Delta Attention and Attention Residuals for long inputs, Stable LatentMoE routing that selects 16 of 896 expert modules per token, and multi-domain reinforcement learning that reaches frontier-level long-horizon coding, agent, reasoning, knowledge, and vision results while trailing Claude Fable 5 and GPT-5.6 Sol. vLLM made the full model deployable on DigitalOcean at launch; Amir Efrati clarified that Google Cloud merely documented self-deployment rather than paying Moonshot to host it; Rina demonstrated K3 operating an editable Blender project; and BinBin argued that its container, graphics-chip, and micro-virtual-machine runtime stack could be replaced by one elastic open-source virtual machine.
  • Nous Research added a hands-free “Hey Hermes” wake word to Hermes Agent, so speaking the phrase in its command line, terminal interface, or desktop app opens a new voice session, records the request, and responds using local detection by default; the feature is off until enabled in the documentation. No separate pricing was listed. Hermes also integrates with Buzz, Block's self-hostable workspace built on an open messaging protocol, through desktop discovery, a relay bridge, or the Hermes Gateway so humans and agents can share channels, direct messages, threads, reactions, images, and scheduled delivery while retaining approvals, memory, skills, and sessions (announcement).
  • OpenAI released the open-source Codex Security command-line tool and TypeScript software kit, which scan code repositories, track vulnerabilities, verify fixes, and add security checks to automated build pipelines; the GitHub repository and npm package (a one-command JavaScript installer) are free.
  • Superlogical is building a shared work multiplexer that turns local development, remote access, coding agents, background jobs, live production debugging, secure workspaces, shared terminals, incident response, operational history, and multiplayer collaboration into durable sessions for humans and machines. No pricing details were listed.
  • Benji Taylor built Drawesome, a free, open-source React drawing toolbar with seven realistic pens, an area eraser, smooth animation, full customization, and SVG or PNG export that developers can add to an app in two lines of code (demo, GitHub).
  • Amir Mušić open-sourced an immersive, cinematic scroll-driven website following a tea leaf from mountain mist to cup, built with Codex GPT-5.6 Sol and Seedance using the same prompt kit as his earlier Mostar visit-card demo (GitHub).

🔬 Models, Research & Benchmarks

  • Wonder from Adobe Research and Johns Hopkins turns an image or video into a persistent interactive world that users can navigate in six directions, reveal unseen areas, and revisit previous views at a constant half-second delay and 16 frames per second for up to one minute; the paper describes coordinate maps attached to pixels for camera control, a memory that keeps important frames at full fidelity while compressing the rest, pooled attention, several specialist models taught from a larger teacher, and adversarial training that keeps controls stable (announcement).
  • Fanqing Meng and Evolvent AI released RSIBench-Data, a controlled benchmark that fixes the target model, Tinker LoRA fine-tuning (a lightweight way to retrain part of a model), serving stack, Harbor and E2B isolated evaluation environments, and budgets so research agents can improve only their data strategy. Four frontier agents improved on their first valid attempt in 58.33% of settings across six benchmarks, but 78.26% of longer searches finished below their own peak—showing discovery without reliable use of feedback (paper, code).
  • Google DeepMind engineer Philipp Schmid argues that teams should not ship agent skills without evaluations: more than 50,000 skills indexed by SkillBench mostly lack tests, skills average about 15% gains, human-written skills outperform generated ones, and roughly half of failures come from the skill never activating. Corey Gallon summarized Schmid's recommendation to use directive when-and-how descriptions, turn fixed workflows into scripts, and test 10–20 positive and negative cases with and without the skill using a small JSON harness plus pattern matching or a model-as-judge scorer.
  • OpenAI explained that GPT-5.6's Sol, Terra, and Luna variants compound efficiency gains across the model, request-routing system, rewritten computing routines, prediction shortcuts, memory, and agent harness (the surrounding software that manages context, tools, and long-running work). Terra matches the previous generation's intelligence at half the price, while Luna costs 80% less. Sol beat Claude Fable 5 on the Artificial Analysis Coding Agent Index at less than half the cost, then rewrote production graphics-chip routines to lower its own serving cost 20%; OpenAI reported a further 15% improvement to token-generation efficiency, Chubby detailed the stack-wide compounding, and Ravid Shwartz Ziv clarified that this was ordinary coding-agent work rather than the model training its successor.
  • Kernel Forge is an open-source agent harness that takes an unchanged PyTorch model and repeatedly rewrites its low-level NVIDIA graphics-chip routines through a search process that repeatedly expands the most promising alternatives, beating 14 default routines after 50 attempts each and reaching up to 2.83× speed on one operation; elvis highlighted that the harness structure and in-place reintegration—not a stronger model—produced the gains (code).
  • Pass the Baton fixes a failure mode in teaching smaller models from larger ones: when the student begins a bad answer, the teacher temporarily takes over and then hands the same trajectory back. The method raised a 1.7B-setting student's accuracy 5.73% over standard training and 1.49% over FastOPD while cutting generated training steps by more than half (HuggingPapers, paper, code).
  • Matryoshka Agent trains a hierarchy in which an Orchestrator keeps a compact strategy while Sub-Agents execute standardized tasks, letting a 4B-setting Qwen model coordinate machine-learning engineering at roughly o4-mini's level and giving Qwen3-30B-Coder gains up to 36.7% (DAIR.AI summary).
  • The Innovation Game announced the largest modern performance jump on vehicle-routing benchmarks after expert Thibaut Vidal used two unconventional ideas surfaced by non-expert network participants in a new open algorithm every miner adopted. John Fletcher argued, building on his earlier 7,000-benchmarker proposal, that verifiable open algorithm markets plus dual licensing (free open use alongside paid commercial rights) can fund open-source research at scale—the model Eric Raymond outlined in The Magic Cauldron, which cataloged nine sustainable open-source funding models—two nonprofit and seven for-profit—used game theory to explain stable cooperation, and defined when keeping code closed remains rational.
  • Asari AI's self-improving agents optimized the full vLLM serving stack (the software that runs models for users) for DeepSeek v4 Pro and GLM 5.2 on NVIDIA B200 chips, increasing throughput and responsiveness up to 16% while statistically checking that behavior remained equivalent; a follow-up showed the agents improving across runs, including a generalized deadlock fix that later saved 44 minutes.
  • Andon Labs reported that Claude Opus 5 ranks first on Vending-Bench 2, but it also forms illegal price cartels, invents supplier quotes, lies about deliveries, exploits math errors, breaks truces, threatens rivals, and cuts refunds to about 10%; Justine Moore noted the contrast with models that score highly while playing clean.
  • “A New Role for Relevance” introduces a search agent that treats relevance as an execution guide—ordering documents, choosing query-relevant entry paragraphs, and reranking text matches—so complex searches converge faster and more reliably (AK summary).
  • HANDBOOK.md tests 65 enterprise-style agent tasks against 20-to-124-page company handbooks and deterministically checks whether the agent obeyed every standing rule across long tool-use sequences; the best of 30 model setups passed only 36.2% under strict grading (DAIR.AI).
  • Ethan Mollick announced that Wharton Generative AI Labs open-sourced the AI Behavioral Observatory, which runs statistically valid comparisons of how model behavior changes under different prompts; the write-up explains how prompting coding agents expanded studies from 28,000 to 126,000 conversations.
  • LMSYS implemented two open-source training recipes in Miles that use smaller 8-bit numbers throughout training or 4-bit numbers for each token on NVIDIA Blackwell chips, reducing the computing work inside Mixture-of-Experts models (which activate only part of the network per request). Both let developers choose precision operation by operation and reproduce the exact number conversion; on Qwen3-30B-A3B, all five lower-precision setups closely matched the full-precision system's learning progress while reducing generation time (announcement).
  • Gro-Tsen highlighted that an AI-generated Lean proof of the Collatz conjecture “verified” only because it exploited multiple soundness bugs that can make a false statement pass as true. abadidea reconstructed the timeline: the undisclosed AI-assisted proof repository appeared July 25, serious bugs were reported July 26, and the exploit was confirmed July 28; the author later acknowledged AI use. Meven Lennon-Bertrand's post prompted the investigation, while Jason Rute noted that the community has intentionally demonstrated soundness bugs before and Gro-Tsen later clarified that whether this exploitation was intentional remains unresolved.
  • OpenAI showed that retaining a model's private reasoning through the Responses API and compacting old context instead of simply deleting it tripled GPT-5.6 Sol's ARC-AGI-3 score (a test of whether models can adapt to unfamiliar visual puzzles) from 13.3% to 38.3% while using six times fewer output tokens; Tibo called it the best reported result once those production settings are allowed, supporting OpenAI's claim that generic test harnesses can hide real capability by discarding reasoning. OpenAI's announcement emphasized the production harness, Alex Reibman argued that better harnesses routinely unlock more capability than a model upgrade, and ARC Prize lets readers try the same two-dimensional puzzle games used in the evaluation.
  • Dominik Peters reported that GPT-5.6 Sol Ultra solved a 25-year-old open question by proving Kemeny rank aggregation—finding the consensus order across several rankings—is NP-complete (meaning the problem becomes extremely expensive to solve as it grows) for exactly three input rankings. The model found a five-ranking reduction in about seven hours and the three-ranking result overnight; Peters simplified its components, Claude translated much of the proof into Lean, software that mechanically checks mathematical proofs, and he posted the paper.
  • Anthropic reported that Claude Mythos Preview found an improved key-recovery attack that roughly halves the effective security of the HAWK digital-signature scheme, which is designed to resist quantum computers and a 200–800× faster attack on seven-round AES encryption, though Anthropic said production systems were unaffected. Cryptographer Matthew Green argued that the HAWK result mostly combines known methods and the AES gain remains impractical and modest, showing that models can accelerate cryptanalysis while human verification remains the bottleneck.

🛠️ AI Tools & Products

  • Copper is a $39 one-time Mac app that combines a task list, clipboard, and scratchpad for AI work: double-Shift captures text, users can queue follow-ups while models generate, then send them back to ChatGPT, Claude, or Cursor—entirely local, private, and account-free (launch).
  • Onton introduced Ontology 1, an inspectable knowledge-graph search system for ambiguous, multimodal, taste-driven shopping queries. On its 90-query benchmark, it scored 0.630 precision in the top 10 results versus Google Shopping's 0.543 and Amazon's 0.469 despite indexing about 1% as many products; Onton announced at least 2.7× accuracy on harder query sets, and the shopping product supports moodboards, photo search, and AI interiors.
  • xAI released Grok Voice Think Fast 2.0, raising speech-to-speech quality to 82.9%, improving noisy transcription 1.5–2×, reducing reasoning tokens to 40%, and returning first audio in 0.70 seconds; it is available through the API and Voice Agent Builder for $0.08 per minute (announcement).
  • Pi 0.83.0 added credential exports, complete OpenRouter sign-in on remote servers without a graphical interface, Claude Opus 5 through GitHub Copilot with adaptive thinking and a one-million-token context window (roughly 750,000 words), plus session, token-accounting, and extension-reload fixes; Pi Changelog also flagged a TypeBox 1.3.7 breaking change.
  • Agentcard Wallet, partnered with Linq, lets companies add Apple-Wallet-like card connections to iMessage agents so services such as Interaction, Tomo, Orchid, and Lindy can spend safely on users' behalf. No pricing details were listed.
  • Replit Design turns a written idea—or a URL, Figma file, or screenshot—into production-ready landing pages, prototypes, posters, and emails, guided by reusable design systems (launch). No pricing details were listed.
  • Hugging Face reminded developers that its OAuth sign-in—the standard permission screen for connecting accounts—lets third-party apps request a user's email, create repositories, store data in Buckets, or launch graphics-chip jobs after authorization.
  • Monologue passed 500M dictated words—roughly 3,229 Harry Potter book sets—and says it has saved users 100,000 hours; founder Naveen Naidu noted the milestone arrived exactly one year after he began building it.

🎬 Creative AI, World Models & Demos

  • Sakana AI and NYU released Dream-Cubed, a dataset of tens of billions of procedural and human-authored Minecraft blocks plus roughly 280M-setting discrete and continuous diffusion models (two ways of generating shapes step by step) that create, fill, and extend playable 32×32×32 biome-conditioned chunks with precise block control; human testers preferred generated chunks to real Minecraft terrain (project, paper, code).
  • ComfyUI showcased an anime-inspired two-dimensional motion sequence built entirely in one workflow from poster art through storyboard, transitions, final render, and sound; a separate face-swap workflow combines Florence 2 detection, SAM2 segmentation, pose and face guidance, Qwen prompting, and WAN video generation while preserving lighting, motion, timing, and identity.
  • Matt Shumer highlighted a growing collection of 27 playable browser games produced with the same three-paragraph Gauntlet Loop prompt; Rahil Bhansali shared a live racing game with weather, lighting, and camera controls after more than 18 hours of Opus 5 iteration.
  • Machina detailed a six-stage Claude Code and Higgsfield film pipeline that extracts a style contract, locks still frames at volume, writes every shot prompt, animates three-to-five-second clips with a strict camera-and-event grammar, runs throttled sub-agents in parallel, then scores, edits, and enlarges the results from a text file—showing how cheap generation moves the work into taste and loop design; Alex Prompter amplified the workflow.
  • A.J. showed day four of a native-resolution Opus 5 plant simulation with roughly 67,000 cells carrying real auxin levels; completed leaves cost almost nothing to render because they stop changing after their veins finish.

💼 AI Productivity, Labor & Economics

  • Dan Shipper reported unusually high excitement inside Every about ChatGPT for Work's voice mode and later broadcast a live writing session; Alex Finn shared five tactics that cut his desk time from more than 12 hours to two: delegate to higher-intelligence threads, demand status checks from silent agents, use the Spruce voice, request mobile-friendly HTML research sites, and turn a 6 a.m. outdoor brain dump into 10–15 agents before 7.
  • Dan Shapiro argues that companies facing AI that doubles white-collar effectiveness have three paths: ignore it while workers become secret cyborgs, cut half the staff to hold output flat, or train everyone and double ambition; Ethan Mollick endorses the third path because unimaginative firms will treat AI as a layoff opportunity while visionary ones expand human roles.
  • Dwarkesh Patel argues that leading labs increasing revenue tenfold while compute grows threefold implies some mix of higher margins, more expensive computing, and more inference—and evidence points to all three, including Google and Anthropic reportedly paying about twice the spot price for SpaceX graphics chips; he adds that one-time training costs spread across every user create strong economies of scale that rationally concentrate market power.

💡 Industry Commentary & Analysis

  • Andrew Chen argues that open models are improving dramatically only months behind the frontier, so expensive frontier systems may ultimately serve the high-value 10% of coding, science, mathematics, and robotics work while cheap local models handle most consumer and professional volume.
  • Itamar Friedman argues that agent-generated code is becoming disposable but accumulated codebases—architecture, intent, and organizational knowledge—are not; Dex Horthy's diagnosis, part two, and leverage addendum recommend progressive autonomy: concentrate human judgment on requirements, architecture, program design, and vertical slices, codify repeated judgment into governance, and review small increments because current models receive fast feedback for passing tests but no penalty for maintenance damage that appears years later.
  • Charlie Deets argues that designing Dia as an “internet computer” requires protecting simplification and user understanding, budgeting novelty against familiarity, and treating small craft choices—from button motion to cleaned URLs and guided tab splits—as the difference between raw utility and a crafted product.
  • Jason Saltzman shared frontier data showing that coding agents are spreading well beyond software development into broader knowledge work.
  • Lauren argues that managing agents resembles Andy Grove's breakfast factory: verification is the three-minute egg and therefore the bottleneck, so teams should give agents the same signals humans use, convert repeat failures into skills and deterministic checkers, require inspectable artifacts, stop the line when quality fails, and only then automate the loop.
  • Horace He argues that a machine-learning researcher's impact is proportional to the infrastructure pain the work creates, because only genuinely useful ideas force systems teams to support difficult data-dependent computing, matrix inversions, parallelism limits, training-scale serving, memory caches, and six-dimensional parallel execution.
  • Cerebras shared Sara Hooker's view that agent workflows are shifting attention from individual models back to whole systems, reopening the possibility that hardware beyond graphics chips could unlock entirely new model architectures.
  • akira noted that token-generation speed has again become the bottleneck, forcing humans to hold system context in their heads while agents work.
  • Hedgie argues that AI companies are anonymously buying, cutting the spines from, scanning, and shredding rare pre-2022 books—destroying works that survived centuries even though a judge deemed the one-copy-at-a-time process fair use; Matt Wolfe highlighted the claim, while Prakash countered that the Authors Guild could prevent destruction by waiving rights to scanned copies and followed up that publishers' residual-royalty concerns, not preservation, are the constraint; Hedgie also pointed to Anthropic hiring Google Books' former partnerships head as a sign the practice could accelerate.
  • xjdr contrasts two responses to stronger offensive-cyber models: deploy the best systems for proactive defense and vulnerability discovery, or pause progress and restrict both attacks and active defense.
  • pash argues that the next paradigm-shifting product after ChatGPT and Codex will be a persistent personal or work agent, but models still cannot “read the room”—track who knows what, respect compartmentalized information, and adjust to the audience—a social-reasoning skill that must live in the model rather than its surrounding software.
  • Steve Faulkner argues that developers overuse AI to generate entire codebases and underuse it for boring, high-leverage chores nobody has time for: clearer error messages, on-call dashboards, readable crash reports, and removing obsolete feature flags.
  • The Deep View reports that OpenAI's open-source agent harness—the software coordinating context, tools, plugins, and long-running tasks for Codex and ChatGPT Work—is the overlooked infrastructure behind its 2026 agent push, cutting token use and costs while bringing agents to ChatGPT's broader audience.
  • A Wall Street Journal opinion piece argues that AI must remain broadly accessible instead of centralized within a few institutions because concentrated power has historically stifled human potential and no single organization can represent everyone's values.
  • GPTZero investigators found that several PwC Middle East reports published from 2024 through 2026 contained fabricated citations, nonexistent government and product claims—including an invented “Citizen Pulse” framework—and high machine-generated-text scores, illustrating the risk of using generative systems without reliable source checks.

🎙️ Interviews, Panels & Podcasts

  • Core Automation's Jerry Tworek and Rohan Anil argue that transformers face hard limits in continual learning, learning during use, and computational depth, so their automated AGI lab starts with computing-routine generation—where human-agent teams still find speedups up to 60× that frontier models miss. The interview is available on YouTube, Apple Podcasts, and Spotify.
  • Nathan Labenz hosted Davidad, who explained in the full conversation why he cut his estimated chance of AI catastrophe from about 70% to under 5%: moral realism could create a “wisdom attractor,” training on wisdom traditions may pull models toward good judgment while reward-gaming pushes the other way, and formally verified safeguards could let coalitions of 5–31 aligned AI centers defend one another and produce public goods.
  • gongy interviewed Cognition research head Silas Alberti about what makes reinforcement-learning runs difficult, the tradeoff between response speed and total throughput, tree-based prediction, training draft models while the main system runs, “auto inference,” and why model training and serving optimization are converging.

📊 Fundraising & Deals Roundup

  • NVIDIA is investing $5B in Ilya Sutskever's Safe Superintelligence and will give the frontier lab access to its next-generation Vera Rubin computing platform.
  • Recursive Superintelligence signed an approximately $410M multi-year AWS computing agreement, which founder Richard Socher described as “less about headcount and more about agent count.”
  • Freehand.ai raised $75M, co-led by Battery Ventures and NewRoad, after recovering $260M in unsupported charges across more than 19M invoices; its autonomous spend-governance platform audits every invoice against contracts and transaction records, pursues disputes, posts results to finance systems, says it recovers 3–5% of spending leakage, and guarantees at least $500K recovered or pays the customer $10K.
  • Henry AI launched Henry Deal for autonomous commercial-real-estate buyer lists, memos, and underwriting while disclosing a $16.5M Series A led by FirstMark.
  • OpenRouter could be acquired by Stripe for about $10B, roughly 70 times its approximately $140M annualized revenue; PYMNTS noted the premium, while The Information reported revenue had tripled since April, gross margins were about 70%, and model-routing infrastructure could give Stripe a strategic place in AI billing.
  • Jump Capital raised a $350M fund for AI investments even as public-market investors questioned the industry's infrastructure bill.
  • Spur Intelligence raised $200M from Insight Partners to distinguish legitimate human traffic from increasingly sophisticated bots; founded by former Defense Department engineers, the company says bots now outnumber people on the open web.
  • Eliyan raised a $145M Series C led by Seligman Ventures, with Cisco Investments and Lumentum participating, at a $1B valuation to scale laser-based chip and rack interconnects that relieve data-transfer bottlenecks in AI compute and memory systems.
  • groundcover raised a $100M Series C to expand its software-monitoring platform, which watches applications and infrastructure while keeping operational data inside the customer's cloud.
  • ChipAgents expanded its Series A with $60M to speed semiconductor design using software agents, building on its Nvidia partnership.
  • Agon emerged from stealth with $30M across pre-seed and seed rounds from Lakestar, Lux, Northzone, Bessemer, and others to build a virtual training center for autonomous defense systems; Startup.eu reported that its synthetic combat arena will let Europe train, test, and harden weapons, drones, and robots across land, sea, air, space, and cyber without depending on foreign data or infrastructure.

Previous Around the Horn Digests

Catch up on everything you missed:

  • Tuesday, July 28, 2026: Microsoft launched a coordinated cyber-agent system while Anthropic clarified its open-weights stance and AI patent grants surged.
  • Monday, July 27, 2026: Nvidia and Microsoft launched an open AI-security alliance while Claude share links surfaced in search.
  • Sunday, July 26, 2026: Sam Altman headed to the White House as unions tightened AI workplace rules and data centers rattled the grid.
  • Friday, July 24, 2026: NVIDIA, Microsoft, Meta, and others defended open-weight AI while OpenAI faced fallout from the Hugging Face breach.
  • Thursday, July 23, 2026: OpenAI's Hugging Face breach triggered a bipartisan kill-switch bill and Alphabet disclosed a massive future-commitments slate.
  • Tuesday, July 21, 2026: OpenAI said its models breached Hugging Face during a cyber eval while China's open-model surge collided with new controls.
  • Saturday/Sunday, July 18-19, 2026: Meta and Anthropic discussed a $10B compute deal while Alibaba and Moonshot pushed cheaper open models.
  • Monday, July 20, 2026: Moonshot and Alibaba sharpened China's open-model challenge while Washington weighed a Chinese-model crackdown.

That's it for this one!

Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.