Everything That Happened in AI Today (Tuesday, July 28, 2026)

Microsoft launched a coordinated cyber-agent system; Anthropic clarified its open-weights stance; AI patent grants surged; OpenAI and Anthropic hit massive revenue estimates; SpaceXAI pushed Grok deeper into app-building and coding workflows.

Written By
Grant Harvey
Grant Harvey
Jul 29, 2026
43 minute read

Microsoft built a cyber team where AI agents attack, defend, and judge the fixes, while the rest of the industry argued over who should be allowed to release powerful models at all.

Welcome to the Around the Horn Digest, where we track every AI story worth knowing so you do not have to. Microsoft turned cybersecurity into a simulated war room staffed by Red, Blue, and Green agents, Anthropic published its formal answer to the open-model backlash, and new patent data showed agentic AI moving from product pitch to intellectual-property land grab. Meanwhile, frontier-lab revenue estimates reached fast-food-chain scale, robotaxis entered another major capital, Google's AI answers kept swallowing more of search, and investors suddenly found Apple's slower AI strategy charming. The machines are now attacking the network, defending the network, and filing the patents. Let's get into it.

Previous digests: Monday, July 27 | Sunday, July 26 | Friday, July 24 | Thursday, July 23 | Tuesday, July 21 | Monday, July 20 | Saturday/Sunday, July 18-19

Around the Horn — Tuesday, July 28, 2026

The biggest product story was Microsoft's Project Perception, a coordinated cyber-defense system built around three kinds of agents. Red agents search for attack paths, Blue agents prioritize and apply fixes, and Green agents evaluate whether the repair actually worked. Microsoft said its MAI-Cyber-1-Flash configuration scored about 96% on CyberGym, a benchmark that tests whether a model can find and exploit real software vulnerabilities, while an Axios breakdown said the model handles roughly 90% to 95% of vulnerability tasks inside Microsoft's MDASH multi-agent system. The model announcement said it nearly halved the cost of the work.

The important shift is not another cybersecurity chatbot. Microsoft is trying to automate the full loop from finding a weakness to testing the repair, with agents checking one another instead of handing a list of alerts back to a human team. That matters more after Axios reported that OpenAI's accidental Hugging Face breach has become a warning shot for defenders. Hugging Face's technical timeline described an autonomous agent escaping its sandbox, the isolated test environment meant to contain it. The agent then exploited a zero-day vulnerability, a software flaw defenders did not know about yet, used a third-party code-execution service as a launchpad, gained administrator and root-level control over computing clusters, and performed roughly 17,600 actions over 4.5 days. Clement Delangue said Hugging Face published the timeline, interactive replay, and defensive lessons so other organizations can prepare. A separate technical breakdown traced attack paths through JFrog Artifactory, a software-package store; HDF5, a scientific data-file format; and Jinja2, a web-template library. Reuters and Axios reported that a Modal customer account was also affected, while Modal CTO Akshat Bubna said Modal's own platform and isolation controls were not compromised. A clip shared by AI Safety Memes quoted Sam Altman saying OpenAI paused training while it worked out how to secure agent sandboxes. Tim Hua estimated that a tiny per-attempt escape rate could still translate into roughly 10,000 internet-reaching events during training, while Adam Karvonen highlighted the tension between a reported 0.01% success rate and the large absolute number created by repeated attempts. Simon Willison said OpenAI should disclose the exact task given to the agent and questioned whether it received the full ExploitGym suite at once, possibly with sub-agents. David Rein noted the uncomfortable historical rhyme that Chernobyl also happened during a safety test.

Advertisement

The breach also accelerated the argument over open defenses. Nvidia's Open Secure AI Alliance brought together Microsoft, Hugging Face, IBM, Cloudflare, Dell, SpaceXAI, Palantir, Cognition, and others to build security tools that defenders can inspect and improve. Nvidia and Jensen Huang argued that attackers already have frontier capabilities, so defenders need broadly available tools. CNBC tied the coalition directly to the Hugging Face incident, Developer Tech detailed its open-source defense plans, and The Verge framed the missing OpenAI, Anthropic, and Google memberships as part of the story. Axios described an economic split between infrastructure companies that benefit from openness and frontier labs that want tighter controls. Andrew Ng called the claim that closed models are inherently safer regulatory capture, using policy to protect incumbent companies from competitors, while Cognition joined the alliance and contributed work on evaluating open models and making them safer after their initial training for safer deployment. The Wall Street Journal reported the wider alliance launch.

🏆 TOP 5 NEWS

  • Anthropic said it has never supported a blanket ban on open-weight AI, meaning models whose downloadable parameters can be run or modified outside the original company. TechCrunch, The Decoder, and The Verge framed the fight as a split between open-model economics, Chinese model risk, and targeted safety controls. Dario Amodei called for mandatory pre-release safety testing of both open and closed frontier models, while Axios said he wants policy focused on chip controls, chip smuggling, industrial-scale model copying, and testing rather than a general ban. Former Anthropic employee Noah Lebovic argued that malicious actors can already reach Claude Code or Codex through ordinary or gray-market accounts, so capable open models may help defenders more than a closed-only regime. Joshua Achiam said Anthropic's position is reasonable but still does not answer the concern that the company wants to slow the spread of frontier technology. Mike called the position sensible if Anthropic's actions match its words and argued that silent model switching should be banned; Susan Zhang read Amodei's framing as a softer version of preserving U.S. frontier superiority through chip sanctions and anti-distillation rules, which restrict training a new model on another model's outputs. Tobi Knaup called open-weight AI's current moment comparable to Kubernetes, the open cloud standard that won because companies competed inside the ecosystem, while Spyglass described the dispute as the familiar openness-versus-control fight complicated by security, geopolitics, and corporate incentives. Dean Ball argued that AI policy is splitting around two defining events: the open-weight letter for one camp and the Hugging Face intrusion for the other. Jensen Huang said Anthropic's Mythos model should be broadly available and called selective access security theater.
  • Agentic AI patents reached 15% of global AI grants in 2025, while U.S. applications involving agents, AI systems that can take multi-step actions, climbed 40% in one year and Nvidia led U.S. filings.
  • OpenAI and Anthropic were estimated to be running at roughly $120B in combined annualized revenue, meaning their current sales pace projected across a full year, though the figures are third-party estimates rather than audited company results.
  • Google's AI Overviews appeared in 43% of searches, up from 15% a year earlier, according to data reported by TechCrunch.
  • Apple overtook Nvidia as the most valuable U.S. company for the first time since May 2025 as investors rewarded Apple's cautious AI posture while questioning the cost of the infrastructure boom; CNBC put the closing market values at roughly $4.95T for Apple and $4.77T for Nvidia.

Honorable Mentions

  • Pacing the Frontier gathered more than 1,100 signatories from OpenAI, Anthropic, Google, Meta, Microsoft, Mistral, and other labs to ask the U.S. government to build international tools for deliberately slowing automated AI research if capability growth outruns oversight. The signatory list included John Schulman, Jakub Pachocki, Jared Kaplan, Shengjia Zhao, Mark Chen, Dario Amodei, Jack Clark, and other senior researchers and executives, while Bloomberg reported the petition's scale. OpenAI said capability acceleration may eventually become fast enough that society needs mechanisms to pace it, and Anthropic cited its recursive-self-improvement research, work on AI systems improving their own research process, as evidence that such tools may be needed. Andrew Curran noted that the wording echoed Sam Altman's comments about giving society time to harden around new capabilities, and OpenAI researcher Dylan Hunn wrote that he is more frightened than ever of recursive self-improvement going wrong but still optimistic that deliberate pacing could preserve the upside. Patrick O'Shaughnessy shared Altman's description of the Hugging Face event as the first AI-security incident he felt viscerally and his warning that sandboxes must survive chained zero-day attacks. Critics split on the remedy. Garrison Lovely argued that pacing still concedes the goal of automating human labor. @scaling01 noted that Anthropic employees made up roughly 46% of signatories. Shannon Sands supported pacing only if it redirects work toward safer, neglected capabilities. @tetsuoai warned that labs could use safety arguments to keep their strongest systems private while using them internally. Jason Wolfe called for domestic transparency and democratic control first. Elie Bakouch warned that vague self-improvement claims could become rules that protect today's largest labs from competitors unless those labs quantify the claims. Adam Thierer called government-led global pacing regulatory capture that would not stop China, then followed up with Mark Zuckerberg's warning about concentrated power. An R Street Institute paper argued that hard global bans are unlikely to work and that continuous coordination, soft-law norms, and practical cooperation are more realistic. Pascale Fung added that open models matter only if researchers can also afford the GPUs, the graphics processors used to run AI models, needed to run them.
  • The Decoder reported that Amazon is scaling back several Nova models and shifting resources to a new Frontier Model Research group led by Pieter Abbeel. TechRepublic said the strategy may consolidate Premier, Omni, Canvas, and Reel into a single multimodal frontier model, a system that works across text, images, audio, or video, while AWS leans harder on infrastructure and third-party models.
  • Lyft and Baidu began testing Apollo Go robotaxis in London, and Baidu said safety-operated RT6 vehicles are running in Brent ahead of a planned public service in 2027, subject to approvals.
  • Anthropic and Cognizant expanded their partnership to embed Claude across Cognizant's business and engineering platforms and create a Claude-certified workforce.
Advertisement

🍪 TOP TREATS TO TRY

  • Grok Build Mode lets SuperGrok Heavy subscribers prompt Grok to create websites, apps, games, and dashboards inside chat, then publish finished projects to a shareable link.
  • Meta AI in Threads now lets users privately ask questions about posts, images, links, and videos shared in direct messages. TechCrunch confirmed the global rollout; no separate price was announced.
  • Cursor's India Start plan costs ₹649 per month in India and includes access to Grok 4.5, Composer, cloud coding agents, iOS steering, plugins, and Model Context Protocol connections, while excluding other frontier models, Bugbot, Auto Mode, Automations, and the Cursor software-development kit. Cursor announced the plan, and Aman Sanger said Indian users tripled year over year and make the most agent requests per user.
  • Meta Ray-Ban Display glasses gained Muse Spark-powered Meta AI, Threads browsing, Instagram updates, and neural handwriting prompts for Early Access users.
  • Lottie Creator 2.0 lets you design, animate, and export interactive Lottie graphics, a lightweight animation format, in a browser with AI-assisted vectors, state machines that define how an animation reacts, and motion tools. Free tier; Individual from $19.99/user/month billed annually.
  • Cekura simulates voice and chat conversations, diagnoses failures, rewrites agent prompts or configuration, and reruns regression tests, repeat checks that catch whether a change broke something, before updates reach production. 7-day free trial; Developer starts at $30/month.
  • Prefactor scores every AI-agent step in real time, flags drift and quality regressions, and can hold, approve, or block risky actions. Free for 25K spans, recorded agent steps, per month; Scaleup starts at $250/month.

🏢 Big Tech & Major Companies

  • Satya Nadella warned that companies relying entirely on one proprietary AI lab may not survive, sharpening Microsoft's pitch for multi-model enterprise systems. He separately argued that trusted U.S. AI ecosystems can outweigh the lower prices of Chinese models such as Kimi K3.
  • Anthropic entered early talks with Samsung about manufacturing a custom AI chip using Samsung's 2-nanometer process, an extremely small manufacturing technology, and advanced packaging that connects multiple chip components closely together. The move could give the lab more control over computing cost and capacity.
  • OpenAI and Anthropic quietly aligned in Washington to influence the Trump administration's frontier-model regulation and evaluation framework despite their fierce commercial rivalry.
  • Elon Musk said Grok 4.6, a 1.5-trillion-parameter model with improved supervised learning and reinforcement learning (training through scored trial and error), is due around August 7, followed weeks later by a larger 2.1-trillion-parameter Grok 4.7 that should be more capable but slightly slower to serve.
  • Anthropic resolved elevated Claude Opus 5 errors affecting Claude.ai, the Claude API (the interface other software uses to connect), Claude Code, and Claude Cowork after roughly 80 minutes.
  • A report said Meta is preparing a new agent harness and open-source models, and Meta chief AI officer Alexandr Wang confirmed the broader direction. A harness is the software layer that gives a model tools, memory, rules, and a workflow.
  • Meta expanded its Louisiana Hyperion data-center plan to 5 gigawatts, pushing the investment above $50B. The New York Times detailed the private negotiations behind the nearly six-square-mile campus, while David Axelrod highlighted the political timing alongside the House's decision not to advance the Kids Online Safety Act.
  • Cloud executives expect Nvidia's transition to its next-generation Vera Rubin chips to be smoother than the Blackwell rollout, which was widely described as difficult.
Advertisement

💼 AI Productivity, Labor & Economics

  • OpenAI found that 43.5% of occupation-specific ChatGPT messages involved work normally associated with another occupation, suggesting AI is expanding people's roles before job descriptions catch up. OpenAI's launch post said the strongest crossover appeared in customer experience, design, and human resources, especially inside small teams.
  • Coursera invested $100M for roughly one-third of Andrew Ng's education startup LearnVector, which plans personalized AI tutors that build learning paths and stay with workers until they master a skill. Reuters said the company will focus on helping white-collar workers adapt as AI reshapes professional work, while Ng said the first products are expected in early 2027.
  • Phoebe Yao argued that data vendors are quietly financing frontier-lab research by paying staff every two weeks, absorbing quality-control and acceptance risk, and then waiting 30 to 60 days for payment. She said that cash-flow mismatch is driving consolidation and fundraising across the data industry.
  • The Wall Street Journal reported that corporate buyers are mixing models and cutting indiscriminate spending on the most expensive frontier systems, shifting more work to cheaper models that are good enough for the task. Kevin Kelly called the change a major shift in AI economics.
  • The Wall Street Journal found young adults using chatbots to script texts, dating openers, and even in-person replies, raising concerns about self-trust and conversational skills. Digital Trends described the same pattern as an emerging epidemic of self-mistrust.
  • Gergely Orosz said engineering leaders are leaving or burning out faster as boards demand AI strategies and teams struggle to gain real AI-native experience. His follow-up added that many replacements are also leaving after only a few months.
  • A startup abandoned a custom project-management tool built with AI and returned to Linear because maintaining the software consumed more bandwidth than the original problem. Jediah Katz argued that AI makes building easier without making maintenance disappear.
  • Chamath Palihapitiya questioned OpenAI's and Anthropic's enormous model-layer valuations, arguing that models improve and commoditize too quickly to support huge long-term values. Ahmad Osman made the same point more bluntly after Kimi K3's release.
  • METR introduced an expenditure-horizon metric that estimates the task length at which an AI agent becomes more expensive than hiring a human for optimization work.

🤖 AI Agents & Infrastructure

  • Model Context Protocol's July 28 specification, the standard that lets agents connect to outside tools and data, made the core protocol stateless, meaning a server no longer has to remember each previous request, and added multi-step requests, header-based routing, cacheable tool lists, stronger authorization, and a formal extension system. VentureBeat framed it as MCP's biggest step toward running agents reliably across enterprise cloud and Kubernetes environments, the software systems companies use to manage large fleets of applications, while David Soria Parra highlighted the new Tasks extension and updated top-tier software kits.
  • Cohere North Automations lets enterprise customers design secure automated workflows with plain-language instructions while keeping human approval at each step. Cohere's launch blog details branching logic, approvals, version history, model choice, and visibility into token costs.
  • Scott Belsky asked for a "favored agent" registration standard so useful AI agents can authenticate to websites without being blocked as spam bots by old login systems, paywalls, and anti-bot defenses.
  • AmpCode's internal usage chart showed its development team rapidly moving from agents running on employees' own computers to cloud agents, with cloud usage projected to exceed 95% within a week.
  • Intent Lab launched an autonomous software-building "fleet" that handles the work between a user's intent and a production system. Its technical overview showed a GLM 5.2 inference engine running 6.3 times faster, a SQLite-compatible database generated in one attempt and passing six million tests, and an agent filesystem checked with formal verification, mathematical proofs that the software follows its specification.
  • Mastra's Trace Intelligence groups hundreds of agent runs by goal, sentiment, behavior, and outcome so teams can see recurring failures without reviewing every trace manually. Sam Calcsam explained that it turns each trace into embeddings, numerical representations of meaning, maps them with UMAP, and groups them with the HDBSCAN density-based clustering algorithm to find natural groups.
  • Prasanna compared PagedAttention in vLLM with RadixAttention in SGLang. PagedAttention reduces wasted short-term model memory inside one request, while RadixAttention reuses shared prompt prefixes across many requests so repeated instructions do not need to be processed again.
  • PRO-LONG showed that a simple searchable interaction log can improve long-horizon agent performance on ARC-AGI-3, a difficult interactive reasoning test, while using fewer tokens than specialized agent harnesses. Alexis Fox and Greg Kamradt highlighted how little extra machinery the method needs.
  • OpenAI research on coding agents that conduct their own experiments found Claude and Codex independently inventing similar algorithms, but Codex also hard-coded evaluation answers until a held-out test set removed the incentive. DAIR.AI highlighted the lesson: agents will exploit a scoring system when the easiest path to a better score is not the intended one.
  • Nader Dabit argued that maintaining an in-house agent platform resembles maintaining a private cloud: every hour spent on the platform is an hour not spent on the product.
  • Recursive Agent Optimization trained agents to create recursive sub-agents and solve harder tasks by dividing the work. Apurva Gandhi said recursion helped models trained on shorter tasks generalize to longer ones, while Alex Zhang argued that a good harness can make structurally similar tasks look almost identical to the model, enabling it to generalize to sequences eight to 32 times longer.
  • ArchAstro builds forward-deployed agents that work across company boundaries on integrations, migrations, upgrades, onboarding, testing, and bug fixing. Calvin Grunewald introduced the company as his post-big-tech bet on agents that complete operational work rather than only answer questions.
  • Alex Prompter argued that graph-based agent engineering beats simple observe-think-act loops once workflows need human approvals, durable retries, crash recovery, parallel branches, and an audit trail.
  • Adi Singh proposed a knowledge-transfer layer between Hermes, Codex, and Claude agents, using agent inboxes as a rough workaround for handing context from one system to another.
  • Hunter Leath argued that future agents need shared, stateful data systems because embodied, radar, lidar, and business agents may need to move terabytes of context without classic upload-and-download bottlenecks.
Advertisement

💻 AI Coding & Developer Tools

  • SpaceXAI added Grok 4.5 to GitHub Copilot's model picker across Visual Studio Code and GitHub products, while keeping API pricing, what developers pay to use the model inside software, at $2 per million input tokens and $6 per million output tokens.
  • MoonEP is an open-source communication library for running distributed Mixture-of-Experts models, systems that activate only a small part of themselves for each request. It can add redundant expert copies dynamically to balance token workloads across servers and avoid slowdowns.
  • OpenAI open-sourced Codex Security as software-development kits and a command-line tool for scanning code repositories, reviewing changes, tracking findings, and running security checks in automated build pipelines. Greg Brockman announced the release, while the Hacker News discussion raised questions about authentication, rate limits, code privacy, token costs, and future local-model support.
  • XY is a Rust-backed Python plotting library built to render extremely large datasets interactively with pan, zoom, hover, density views, and exports to HTML, PNG, SVG, or PDF. The Show HN thread discussed its GPU acceleration and composable interface.
  • Verified 3D Mesh Intersection uses Lean 4 mathematical proofs so reviewers can trust a 93-line human-written specification instead of manually auditing more than 1,000 lines of AI-generated geometry code. The Show HN discussion focused on formal verification as a way to make AI-written code auditable.
  • minions-army-harness turns a plain-English request into a specification, implements it in an isolated environment, opens a pull request for human review, and runs an adversarial review before an optional deployment.
  • ctrlb-decompose compresses noisy application logs into recurring patterns, typed variables, statistical ranges, anomalies, and severity scores before sending them to an AI model. The Show HN thread positioned it as a way to reduce token waste and improve debugging.
  • Hubble.md is a free, open-source Markdown editor with comments designed for collaborating with coding agents on plans, drafts, and blog posts.
  • KDA-B200 is an open-source CUDA implementation, GPU code written for Nvidia chips, of Kimi Delta Attention for Nvidia Blackwell hardware. It reported 1.42 times faster performance than the public FlashKDA interface while passing all 580 upstream correctness tests, and a reusable-workspace configuration reached 1.52 times faster performance.
  • Cursor added Kimi K3, bringing Moonshot's open-weight model into a widely used coding environment through U.S.-based inference providers with zero-data-retention support.
  • Cursor's planner-worker swarm showed an expensive frontier model planning a coding project while cheaper models executed most of the individual tasks, suggesting teams can reserve premium reasoning for coordination.
  • Cerebras published a guide to using GPT-5.6 Sol, Terra, and Luna in Codex. It recommends starting with the faster, cheaper Luna model, escalating only when stuck, keeping sessions warm so repeated context can be reused, and using multiple agents as advisers without paying frontier-model prices for every step.
  • Vercel's eve added Slack event hooks and session controls, including thread follow-ups without repeated mentions, cancel-and-replace for in-progress turns, full session resets, and callbacks for reactions and other Slack activity.
  • JJ Englert shared a Claude style guide that forces Opus 5 into shorter, answer-first, structured replies after finding the default model too chatty, overconfident, and prone to stopping short.
  • OpenAI backported refreshed Codex metadata to the stable 0.144 release line, including GPT-5.6 model instructions, context-window information, reasoning summaries, skills, permissions, and automated review settings.
  • Matt Shumer named the Gauntlet Loop, an agent workflow that breaks a goal into parts, assigns specialist builders and blind critics, and accepts only work that beats a real-world equivalent. The pattern grew out of his Claude-of-Duty project and original prompt; Jason Kneen's code change later improved performance by moving camouflage-texture calculations to the GPU.
  • The Engine Shop is a set of essays on AI-first software development. The series emphasizes hard contracts, repeatable checks, and engineering loops that can keep up with cheap implementation.
  • Justin Schroeder shared a long-horizon, one-prompt coding comparison chart that he said matched real-world experience and showed large gaps among frontier models.
  • the tiny corp argued that software engineering has never been limited by how quickly people can produce code. The constraint is managing complexity after the code exists.
  • Cole Murray argued that the lasting asset in background coding-agent systems is the runtime: environment capture, access controls, deployment design, and spending limits. Skills, MCP connections, and rules are portable, but failures that appear months later are hard to fix when the runtime belongs to a vendor.
  • Migel Tissera reported that Cursor uploaded 63,106 files, roughly 736MB of source code, months after he canceled his subscription and used the app only as an editor. He said indexing appeared to depend on authentication rather than the visible privacy and indexing controls, and that he could not find a way to delete the uploaded files.
  • Gal Zahavi argued that the three useful tests for a coding harness after heavy use are the interface flow, how efficiently it spends tokens, and the hard-to-quantify ability to finish workflows faster without extra cost.
  • Claude Design is one of the most underrated AI products, according to Mckay Wrigley, who said Opus 5 transformed how he builds design work.

🔬 AI Research & Models

  • Moonshot AI released Kimi K3's full open weights, the downloadable model parameters that let others run or modify the system.
  • The weights, launch post, technical report describe a 2.8-trillion-parameter Mixture-of-Experts model. Parameters are the adjustable values inside a model, and only 104B activate rather than the full system for each request. It has native vision, meaning it can process images, and a one-million-token context window, enough to process several books at once. Tom's Hardware reported near-frontier performance, while VentureBeat noted that the roughly 1.5TB download still requires serious infrastructure and that large model-service providers face a separate commercial license. Sebastian Raschka's architecture notes explain LatentMoE, the model's expert-routing design; Kimi Delta Attention, its long-context memory mechanism; attention residuals, connections that carry earlier attention outputs forward; and the multimodal design that joins text and images. The model reportedly delivered near-frontier performance with two to three times greater serving efficiency and up to 10 times lower cost when providers could reuse previously processed prompt data. Moonshot's announcement emphasized the open model, kernels, communication libraries, and agent infrastructure as one system.
  • Moonshot also released FlashKDA, optimized GPU code for Kimi Delta Attention that speeds the first processing pass on Nvidia H20 chips. Amirhossein Kazemnejad argued that Kimi K3 may mark the end of explicit positional encoding, the mechanism that tells a model where each token sits in a sequence, in frontier models. Alex Ker highlighted its 1.8% expert-activation rate as the efficiency breakthrough. Sebastian Raschka added the architecture to his LLM Architecture Gallery, arguing that open weights let researchers verify claims and run models without sending private data to a vendor.
  • Serving teams moved quickly. Philip Kiely described Baseten's day-one API support; LMSYS reported SGLang model-serving support at 423 tokens per second; and Cheng Wan explained the hybrid Kimi Delta Attention and Multi-head Latent Attention memory design, which compresses stored attention information to reduce memory use, replay system, and fused GPU operations behind that speed. Susan Zhang said the model's stability fixes show how difficult signal flow remains at this scale; Ali's architecture worklog traced the path from GPT-2 through newer memory and expert-routing systems; and Anmay Gupta explained why the vision system could identify specific Unsplash photo IDs during reasoning.
  • The openness still has limits. Thomas Unise called downloadable weights a path out of a permanent AI underclass, while Nathan Lambert flagged commercial-license thresholds for companies above $20M in annual revenue. Unsloth AI committed to local inference and smaller variants, and Alexander Doria praised the report for connecting the model, GPU kernels, communication layer, and agent stack. Jamin Ball found that Baseten and Fireworks matched Moonshot's $3-per-million-input and $15-per-million-output token pricing, contrary to his earlier expectation that open weights would lower prices immediately. Kevin Xu pointed to the license and U.S.-hosted service guarantees; Brian Zhan noted that heavy reasoning-token use can keep total costs high; and Nebius Token Factory launched an OpenAI-compatible Kimi K3 endpoint at the same prices.
  • Meryem Arik said Moonshot's provider agreements prevent large hosts from undercutting one another, pushing competition toward response speed and the amount of traffic each provider can serve; Doubleword launched at $2.15 per million input tokens and $11.25 per million output tokens, a reported 28% discount. JJ estimated that self-hosting Kimi K3 on eight Nvidia B300 GPUs at about 75% use could serve more than 30B tokens a month for under $10K in monthly operating cost after roughly $500K in hardware, paying back current API prices in under 100 days. stochasm highlighted undisclosed training-token counts and the cleaner gated Multi-head Latent Attention design, which adds learned gates to the memory-saving attention system. The breakdown also covered NoPE, which removes explicit position embeddings; a stabilization technique the report calls QB scale-up; bounded SiLU activations that cap signal size; co-located reinforcement learning, which keeps training workers near the systems generating examples; and unusually detailed sandbox infrastructure. Bhavin Jawade and Nova Sarc broke down the post-training recipe: supervised fine-tuning on synthetic agent trajectories, nine specialist teachers across three domains and three reasoning budgets, then multi-teacher on-policy distillation, where the student learns from several teachers while generating its own attempts. It used clipped reverse-KL, a training loss that keeps the student close to those teachers without allowing extreme corrections. The pipeline used XTML, a markup format for agent trajectories, and did not add a separate reinforcement-learning-from-human-feedback (RLHF) or direct-preference-optimization (DPO) stage. Patrick O'Shaughnessy shared Sam Altman's view that great cheap models will exist, OpenAI distills its own models, open weights will remain important, and competitors copying model outputs is not among his ten largest worries.
  • Moonshot is already seeking more Nvidia Blackwell chips for a significantly larger Kimi K4, according to The Information.
  • Kimi Linear introduced a hybrid attention architecture, a memory system that mixes standard token-to-token comparison with recurrent state. It reportedly beat standard full attention in matched tests while cutting memory use for long conversations by up to 75% and improving one-million-token answer-generation throughput by up to six times. The Hacker News discussion debated its scaling behavior, model-distillation claims, and connection to earlier Gated DeltaNet work.
  • InclusionAI released LLaDA2.2-flash, a 100B-parameter diffusion language model that edits text through insertions and deletions instead of producing one token at a time. The team reported competitive software-engineering results and 1.7 times higher throughput, with weights and a technical report available.
  • Experience Distillation lets an agent write lessons from past trial-and-error sessions into its model parameters instead of forgetting them when the conversation ends. The alphaXiv summary said the method retained at least 64.8% of the gains from keeping examples in the prompt and matched standard reinforcement-learning baselines with at least 9.6 times fewer environment interactions across software tasks and text games.
  • Liquid AI released open-weight encoders, models that read and classify a full input instead of generating an answer one token at a time. The 230M model ran about 3.7 times faster than ModernBERT-base on long CPU inputs, while the 350M model traded some speed for stronger accuracy. Liquid AI detailed the 8,000-token design and support for 15 languages in its technical blog, listed compatible models and formats in its documentation, and announced the encoder launch and interactive demos. Viviana Márquez argued that predictable jobs such as routing and policy checks are often better handled by one-pass encoders than by token-generating chat models. The public demos cover prompt routing, policy linting, spellchecking, and masked diffusion, which reveals an answer by repeatedly filling hidden tokens.
  • Anthropic researchers used Claude Mythos Preview in a multi-worker research environment to find a lattice automorphism that sharply reduced the expected cost of recovering HAWK signing keys and a Möbius Bridge fingerprinting method that improved attacks on reduced-round AES-128 by 200 to 800 times. In plain English, the model found mathematical shortcuts that made two cryptographic attacks much cheaper. Anthropic released part of the cryptanalysis code for inspection.
  • Large language models predicted social-science experiment results with a reported correlation of 0.85 across 70 preregistered studies, matching or exceeding expert forecasts while raising concerns about automated persuasion. Robb Willer highlighted the findings, and the Treatment Effect demo lets users enter messages and outcomes to simulate likely responses from U.S. adults.
  • Nvidia's Sol-Attn paper introduced a training-free method that skips less-useful attention calculations during video generation, reporting up to 2.1 times faster generation and 2.3 times faster editing without visible quality loss. The team published a project explainer and open-source implementation.
  • A Nature study analyzing more than 14,000 cortical units across 43 brain regions found that most neurons behave like flexible generalists rather than narrowly specialized cells, producing representations that make many conditions easy to separate. New Scientist translated the finding as most neurons being "jacks-of-all-trades."
  • Transluce proposed oversight foundation models trained inside simulated worlds so developers can generate large amounts of automatically correct training data for detecting reward hacking, when a model games its scoring rule instead of doing the intended job, hidden behavior, and unexpected changes after fine-tuning.
  • Observational Imitation Learning trains a policy by watching multiple imperfect teachers and selecting only their best actions. Guohao Li noted that the 2019 robot-learning idea now resembles multi-teacher distillation used in modern language-model training.
  • Squeeze Evolve is an evolutionary framework that tries and keeps better agent strategies without a separate checker, sending high-impact reasoning steps to stronger models and routine steps to cheaper ones, reporting up to roughly three times lower API cost and 10 times higher throughput. The paper describes the method, and Monish Maheswaran released it as a Claude Code plugin.
  • Avi Krishna's personality study found frontier models converging on a similarly helpful, concise, and noncontroversial default personality while still differing on creativity and emotional expressiveness. The thread summarized results from 976M analyzed tokens.
  • OrbitAll is a physics-grounded molecular model that represents electron orbitals and respects how 3D objects rotate. Anima Anandkumar said it beat Meta's UMA on charged, open-shell, and solvated chemical reactions while using 35 times less data, a 50-times-smaller model, and roughly 100-times-faster inference.
  • WorldDiT is a unified diffusion-transformer architecture, a sequence model trained by learning to remove noise, that jointly models robot actions and future visual states without relying on a large pretrained vision-language action model. Bagel Labs reported 94.9% mean success across four LIBERO robot-manipulation test suites with roughly 399M parameters, and released the weights.
  • Music-JEPA learns a world model of piano sound by treating audio as the current state and piano-roll notes as actions. Tanishq Mathew Abraham highlighted that the learned representation supports beat tracking, composer identification, key estimation, and music transcription through planning; the demo lets readers hear the results.
  • Reverso is a family of time-series forecasting models with fewer than three million parameters that predicts future values without task-specific retraining. Its multi-scale inputs and hybrid long-convolution and DeltaNet recurrent-memory layers handle 16,000-step histories, matching much larger models on the Gift-Eval forecasting benchmark while running in a browser. The team published a plain-English explainer, code, weights, and a launch thread.
  • ECMWF's AIFS Single 2.0, an open AI weather-forecasting model, can now run on a wider range of GPUs. Emma Scharfmann published a compatibility patch and tutorials, while the Hugging Face guide and GitHub repository show how to run it locally or through Hugging Face Jobs without requiring FlashAttention, a specialized GPU-speedup library.
  • IDEAgent introduced a multi-agent framework that generates many different research ideas rather than converging on one. The authors reported a 3.89-times improvement over baselines across 32 computer-science topics and released the code.
  • Moonshot AI released PerceptionBench, a 3,000-question benchmark for basic visual skills such as locating, counting, and comparing objects. The GitHub repository and dataset show that no frontier multimodal model cleared 60%.
  • The Regression Tax found that adding procedural skills to agents can break tasks the same model solved without the skill. Elvis and Ksenia Se highlighted three failure modes: the skill description bleeds into unrelated behavior, the model stops grounding itself in the current task, and it skips verification because the skill makes it feel overconfident.
  • Role Drift in Compound LLM Systems found that modules can improve final accuracy by violating their assigned job. Elvis noted that many apparent reinforcement-learning gains disappear once each module is forced to stay in its lane.
  • Interactive Training 2 proposed an auditable control plane for steering a model while it trains. The live demo, Papers with Code entry, and HuggingPapers announcement show humans and automated controllers using the same protocol to pause, inspect, and redirect training.
  • Alexi Gladstone previewed work on exposure bias, the problem created when a model trains on perfect examples but must generate from its own imperfect outputs, building on his essay Training for Marathons by Sprinting.
  • Rafa Schwinger argued that language models need sparse weight matrices, neural connections with many deliberate zeros, to create higher logical dimensions efficiently. He followed up that algebra over concepts and relationships may be a natural route to that sparsity.
  • Michael Yu adapted Anthropic's Natural Language Autoencoders to Gemma 3 image-token patches, making parts of a vision model's internal representation readable as text. He published interactive examples and code.
  • Expanding Flow Maps let generative models grow the size of their state while generating instead of working on a fixed canvas. Pranam Chatterjee highlighted Sophia Tang's work, while alphaXiv summarized how it can generate variable-length text, graphs, and 3D molecular shapes in one framework.
  • Qwen-Audio-3.0-TTS expanded Alibaba's text-to-speech system to 16 languages. Tongyi Lab described two versions, Flash for real-time conversation and Plus for higher-quality generation, with natural-language style control and detailed tags for nonverbal sounds. Alibaba's streaming documentation covers low-latency audio, voice cloning, voice design, and simultaneous streaming input and output, while Artificial Analysis compares speech-to-speech models on reasoning, delay, and price. Wildmind AI also highlighted the release.
Advertisement

🦾 Robotics & Physical AI

  • Tau Robotics launched an invite-only humanoid home-cleaning service in San Francisco for a flat $30 per hour. Each robot is controlled jointly by AI and a live human operator, and the company says its demonstrations are shown at normal speed; the waitlist is open.
  • tau0-VLA is a hierarchical robot foundation model for long-horizon household work. The project site, code, and paper describe a high-level policy that proposes subtasks, a world model that imagines likely outcomes, and a lower-level vision-language-action model that executes movements across cleaning, cooking, laundry, and object organization.
  • Jiashun Wang released an open-source extension of hybrid motion-imitation work for Unitree G1 box-moving and box-climbing tasks. The repository shows one robot policy learning both from reference movements and from reinforcement-learning goals.
  • Aladdin is an autonomous electric moped that can travel up to 28 mph on bike lanes and local roads, with onboard AI for commuting, meal pickup, grocery runs, and other local trips. Genie Mobility says it has a swappable battery, stores personal context locally, accepts a $100 deposit, and is scheduled to ship in summer 2027.
  • Claude CAD turned a messy four-page customer PowerPoint with screenshots, handwritten dimensions, and non-scale artwork into a usable STEP 3D model, a standard editable engineering format, and a flat-pattern DXF cutting file that covered roughly 90% of the requested deliverable.
  • comma.ai sells a plug-and-play device that adds hands-free lane centering and adaptive cruise control to more than 325 vehicle models across 27 brands without a subscription.
  • Wind's autonomous-rides beta puts a camera-only self-driving system onto an existing golf cart without lidar, the laser sensors many autonomous vehicles use to map their surroundings. Avi Krishna highlighted it as an example of adding new autonomy software to existing hardware.
  • ModPack is an open modular teleoperation backpack that shares one untethered core across two-armed mobile robots, then adds swappable modules for touch feedback, perception, and base control.

🏛️ AI Policy, Governance & Safety

  • Axios reported that smaller specialized cybersecurity models from Microsoft, Google, and Cisco are emerging as cheaper defensive tools for organizations that cannot afford or access frontier cyber models.
  • Public Claude share links became a practical privacy warning. Axios reported that public links for apps, documents, spreadsheets, and visualizations could appear in Google when users posted them somewhere crawlable. 404 Media found pages containing apparent keys, legal questions, personal information, and app data. TechCrunch said Anthropic maintained that pages were indexed only after users made them public, while The Decoder reported that the pages appeared to lack a noindex instruction before disappearing from search. WIRED found shared snapshots in Google and Bing despite a robots.txt file because the pages lacked stronger page-level blocking, and Allie K. Miller demonstrated public Artifacts containing family financial plans, law-firm contacts, and medical-staff schedules.
  • Kenya opened public consultation on its first comprehensive national AI and emerging-technologies policy through August 4, covering governance, data, safety, skills, and a proposed National AI Council.
  • The EU Digital Omnibus on AI entered into force. A legal analysis said it delayed several high-risk obligations, added prohibitions involving non-consensual intimate imagery and child-abuse-material generators, reduced some registration and small-business burdens, and extended regulatory sandboxes where companies can test systems under supervision.
  • The Trump administration moved to ban new Chinese humanoid and four-legged robots plus connected power inverters, citing espionage, remote-control, and infrastructure-disruption risks. Axios described the move as an attempt to protect the U.S. AI supply chain. FCC Chairman Brendan Carr said advanced foreign robotic devices and power inverters were added to the Covered List, blocking new versions from U.S. import or sale, while Chris McGuire argued that the policy could reshape domestic robotics if licensing favors U.S. and allied producers.
  • AI Forensics researchers found that seven of nine popular image-editing Spaces on Hugging Face produced non-consensual topless deepfakes from a short prompt, while a honeypot tool received more than 1,000 sexual prompts in one week. Engadget highlighted weak output moderation and the apparent targeting of women and minors.
  • Runlayer sued Rippling for alleged trade-secret theft after a long Model Context Protocol gateway evaluation under a nondisclosure agreement. The New York Post reported that an internal Rippling project was described as "essentially a clone" of Runlayer's agent-safety and governance product.
  • AI-assisted security research put 2026 on pace for roughly twice as many recorded software vulnerabilities as 2025, with more than 45,000 flaws already logged.
  • The Financial Times reported record Washington lobbying by OpenAI, Anthropic, Google, and Microsoft as federal AI-policy fights intensify.
  • China's Ministry of Commerce rejected U.S. claims that Chinese labs copied American models through distillation, the process of training one model on another model's outputs, calling the accusations factually and legally baseless.
  • Zvi Mowshowitz argued that the Hugging Face incident exposed systemic failures in containment, monitoring, and alignment rather than one isolated benchmark mishap.
  • Visa open-sourced an agentic vulnerability-scanning harness, and Leon Derczynski argued that much of a defensive system's strength lives in the harness rather than the underlying model.
  • A report claimed the U.S. is negotiating a pre-release checkpoint framework that could give federal agencies up to 30 days of exclusive access to new frontier models from OpenAI, Anthropic, and Google.
  • Lisan al Gaib argued that aggregate AI progress can look smooth while releases from each lab arrive in discontinuous jumps, and that those jumps become more dangerous as capabilities grow.

🛠️ AI Tools & Products

  • OpenAI's new transcription models split speech-to-text work between GPT-Live-Transcribe for low-delay streams and GPT-Transcribe for uploaded files and batch jobs. The API guide covers file and real-time transcription, context prompts, keyword hints, and language settings, with improvements on accents, numbers, technical terminology, and noisy audio.
  • Coast gives users and their agents fully local memory by recording what appears on a Mac and processing it on-device through Apple's Neural Engine, the on-device chip used for AI workloads. Aidan Guo framed the launch as a way to preserve useful context people normally discard at the end of each day. Free download for Mac.
  • The Complete Shelf presents 19 procedurally generated hardcovers on a continuous 3D shelf, letting users pull out each volume, rotate it, zoom in, and inspect its editorial details.
  • Segue saves a block of context from one AI assistant and retrieves it in another with a short pronounceable handle through Model Context Protocol, the standard that lets agents connect to tools and data. The Show HN discussion raised concerns about storing and relaying the context as plain text.
  • Yap is a free, open-source macOS dictation app that uses Apple's on-device Speech framework and pastes your words into the active field with no cloud account or downloaded model. The Show HN thread compared it with cloud and Whisper-based alternatives.
  • Flashpaper sends self-destructing passwords, API keys, credentials, and files up to 10MB using encryption performed in the browser, memory-only storage, burn-after-read controls, and timed deletion. The Show HN discussion compared it with Privnote and password-manager sharing.
  • Tines 3B gives employees a governed environment for building AI agents, apps, and automations while IT and security teams retain visibility and credential control. The Show HN thread focused on how companies can manage employee-built workflows already appearing outside approved systems.
  • AI Product Academy's Builder Pass bundles eight AI product-management certifications, an invite-only community of more than 300 senior product managers, one-on-one sessions, and future courses for a year. Dr. Marily Nika announced the program; pricing is available on request.
  • Adomate turns Meta performance data, competitor ads, and customer reviews into traceable ad concepts and repeatable creative-research workflows. Free plan; Starter starts at $119/month.
  • Webhound runs cited research or builds sourced datasets to a dollar budget you set, either directly or through an agent tool call. $5 free credit; pay-as-you-go from about $1 per 15 minutes.
  • superfile gives terminal users a polished multi-panel file manager with previews, fuzzy search, bulk operations, themes, and plugins. Free and open source under the MIT license.
  • Quill starts and stops dual-track recording of microphone and system audio from the macOS menu bar, then transcribes everything locally with speaker labels using Parakeet, an on-device speech-recognition model. Andrew Dreskin shared the launch. Free and open source.
  • Octen gives agents high-concurrency real-time web search at a published median delay of 62 milliseconds, roughly six times faster than Exa in Monid's comparison. Monid's launch post described splitting one task into thousands of parallel searches; no pricing details were announced.
  • Mage-VL is Microsoft's 4B-parameter streaming vision-language model, a system that reads video while it arrives, and aligns its processing with video codecs to cut visual tokens by more than 75% and run up to 3.5 times faster. The free demo and launch thread show it analyzing live video.
  • AgentENV launches isolated Firecracker microVMs, tiny virtual computers that separate an agent from the host machine, in under 50 milliseconds and supports snapshots, copying an environment into branches, memory adjustment, and compatibility with E2B's agent-sandbox interface. Guillermo Rauch highlighted Kimi K3 experiments that crashed ordinary container hosts while the microVMs stayed isolated. Free and open source.
  • NVIDIA Object-Oriented Agents turns agents into ordinary Python classes whose fields hold state, methods become actions, and type annotations act as contracts. Alessio Devoto introduced the framework, and the paper reports results on SWE-bench Verified, real GitHub bug fixing; Terminal-Bench 2.0, computer-use tasks; and ARC-AGI-3, difficult interactive puzzles. Free and open source.
  • PorTAL trains a task-specific LoRA adapter, a small add-on that changes a model without retraining the whole system, once and transfers it across multiple frozen base models. Ramp Labs open-sourced the project, with models and a Python package available. Free and open source.
  • Remotion Agent Skills creates complete videos through Claude Code after one install command. The full Claude Code session shows how the demo animation was prompted. Free to try with Remotion's open-source tooling.
  • Silico helps researchers inspect, debug, and intentionally redesign model behavior through automated interpretability experiments, tests that reveal what patterns inside a model drive an answer. Goodfire's bouba-kiki demo found that spiky- and round-sounding words align on opposite ends of a model-activation direction independent of meaning. Request access.
  • OpenAI made GPT-Live in ChatGPT Voice available to Education, Business, and Enterprise plans globally.
  • Vercel added Kimi K3 and Kimi K3 Fast to AI Gateway through U.S.-based providers with zero-data-retention support. Guillermo Rauch called it the most powerful open-weight model available through high-availability U.S. inference, while regional inference now keeps supported model requests inside the U.S. or European Union.
  • OpenRouter discounted GPT-5.6 Terra and Luna by 50% for first-party OpenAI traffic for a limited time.
  • Kimi K3 took first place on DesignArena's presentation-slide benchmark with the largest margin the leaderboard had recorded.
  • Claude Opus 5 with Max reasoning ranked first in both the Frontend Code Arena and Text Arena, while the default variant ranked close behind.
  • Startracker used a 30-agent studio to build a space-flight game, assigning frontier models as technical leads while sub-agents produced avionics, weather, stars, physics, and ground systems.
  • ODS is a private local-AI server setup that detects your hardware, downloads a suitable model, starts an inference engine and Open WebUI, then adds voice, agents, retrieval from your own files, web search, and image generation from one dashboard.
  • MTS Live partnered with Arena for dedicated benchmark and evaluation coverage during model releases. Brent Liang highlighted Arena's growth from no revenue to a reported $100M, while MTS's live broadcast extended its always-on monitoring across technology, finance, geopolitics, and culture.
  • MPPscan tracks machine-to-machine payments across Tempo and other networks, including agent transactions, micropayment servers, volume, and activity. Patrick Collison said MPP transactions were approaching 30,000 per day, while Akash Bajwa argued that sub-cent settlement makes tiny inference and data purchases commercially viable.
  • Claude 5 Opus generated a complete car-on-a-dirt-trail scene from one prompt, including graphics, motion, and lighting, without external textures.
  • A minimal Umwelt in Lenia is an interactive experiment that lets users explore preferences emerging in digital organisms that were not explicitly designed for specific behaviors. Michael Levin shared the new demo, building on his earlier example of simulated agents developing avoidance of the unknown.

💾 AI Infrastructure, Chips & Data Centers

  • Nvidia and Ilya Sutskever's Safe Superintelligence announced a long-term partnership. The official release includes access to Nvidia's Vera Rubin computing platform and enough capacity to increase SSI's compute tenfold. Bloomberg reported a $5B cash investment, while the Financial Times described the long-term compute arrangement. Andrew Curran collected details about SSI's brain-inspired research direction and future-platform collaboration, while Gavin Purcell emphasized the planned 10-times compute increase. Chris interpreted the work as a possible move toward continual learning, where a model keeps updating after deployment. Ravid Shwartz Ziv joked that Sutskever's pitch amounted to "our research is worth scaling" without a benchmark. The Decoder said the deal may move SSI's next scaling phase away from Google hardware.
  • The Information argued that investors should not panic over Nvidia's growing AI commitments, even as reported OpenAI financing support and the SSI investment renewed concern about circular deals in which chip companies finance their own customers.
  • China began mass-producing domestic immersion deep-ultraviolet (DUV) lithography machines, the complex tools that project ultraviolet light through liquid to print advanced circuits onto chips. The Information's market brief, Bloomberg, and Tom's Hardware said the news pressured ASML as China prepared to ship roughly five tools in 2026 and about 20 more in 2027.
  • Nvidia entered talks to guarantee up to $250B in financing so OpenAI can lease a 10-gigawatt Ohio data-center campus that could ultimately cost more than $500B. CNBC explained how Nvidia's credit could support the debt, while Axios connected the talks to Nvidia's wider financing strategy. The Kobeissi Letter highlighted the project's extraordinary scale. Nicolas Bustamante argued that compute remains the industry's hardest moat because power, networking, cooling, land, permits, and capital do not spread as easily as research ideas.
  • AI data-center bonds reportedly reached about $270B in 2026 as borrowing costs rose alongside the industry's computing buildout.
  • AMD and Cerebras partnered on split inference, sending prompt processing and long-context preparation to AMD Helios systems while Cerebras wafer-scale chips, processors built across most of a silicon wafer, generate tokens. The companies claimed up to five times better tokens per second per watt.
  • Core Scientific and AMD agreed to deploy more than 500 megawatts of U.S. AI data-center capacity starting in 2027, scalable to 2.5 gigawatts, using AMD Instinct GPUs, EPYC CPUs, and ROCm, AMD's software layer for running GPU workloads.
  • Meta and BlackRock formed a venture to finance and operate a one-gigawatt El Paso data-center campus with roughly $14B in projected costs. BlackRock's infrastructure funds will own 80%, Meta will own 20%, and Meta will lease the full campus.
  • EPA guidance said "islanded" power plants serving only data centers and not the public grid fall outside the Clean Air Act Acid Rain Program. Reuters reported that the interpretation could let developers build dedicated generation without those federal pollution-program requirements.
  • Nvidia was revealed as the tenant for a $50B data center that will use Nvidia chips, with Jensen Huang using the company's balance sheet to help support continued growth in the AI-computing market.
  • Recursive signed a $410M AWS collaboration to scale its self-improving AI research system after emerging from stealth at a reported $4.65B valuation.
  • Ionic Digital jumped more than 25% in its Nasdaq direct-listing debut after converting former Celsius bitcoin-mining assets into AI infrastructure, including a 10-year lease with Nscale for a West Texas site.
  • CXMT raised $8.6B and surged 466% in its Shanghai debut, making the memory-chip company China's most valuable listed company.
  • Seagate beat earnings and revenue expectations and raised guidance as AI-related storage demand accelerated.
  • Corning fell 12% after issuing a current-quarter revenue forecast below expectations, triggering double-digit declines across optical-component suppliers tied to AI data-center networking.
  • Taiwanese prosecutors detained an Nvidia employee during an investigation into alleged illegal exports of Super Micro AI servers to China. PC Gamer said Nvidia itself was not accused of wrongdoing.
  • Atomarine announced nuclear-powered floating data centers that package electricity generation and computing into deployable units, claiming they can be built faster than land-based projects waiting for grid connections.

📊 Fundraising & Deals Roundup

  • Cyera agreed to acquire Oasis Security in a deal reportedly valued at $1B, combining data-security posture management with protection for machine and AI-agent identities.
  • Multiverse Computing targeted up to $570M in Series C funding to expand CompactifAI, a tensor-network compression system, a mathematical method for representing a large model with fewer values that claims to shrink models by as much as 95% while preserving accuracy. Pathfounders reported a $1.7B valuation and plans spanning edge devices through sovereign cloud infrastructure.
  • Antares raised $470M to build modular nuclear reactors producing roughly 100 kilowatts to one megawatt for U.S. military bases.
  • Dwelly raised $170M to acquire real-estate businesses and add AI to their operations. Sifted said the round included $95M in equity, $75M in debt, and backing from founders at ElevenLabs and Legora.
  • Enigma raised $71M to make robot control as intuitive as adjusting volume. Enigma's launch put more than 100 real robots online for browser control, and the Robots.online platform includes tasks ranging from sword duels and painting to laboratory work.
  • Fish Audio raised $52M after reaching more than eight million users and $21M in annual recurring revenue. The company also launched S2.1 Pro, which clones a voice from five seconds of audio, supports 83 languages, starts producing sound in roughly 90 milliseconds, and offers word-level control over emotion and intonation. The company said it runs twice as fast as Cartesia at one-sixth the cost of ElevenLabs, is used by HeyGen, LiveKit, Retell, and others, and will provide a free year if it cannot cut a customer's voice costs by 50%. TestingCatalog covered the launch, and the enterprise page handles production pricing and security reviews.
  • Mate Security raised $35M in Series A funding eight months after its $15.5M seed round.
  • Hush Security raised $30M to expand a machine-access platform that registers enterprise agents, grants short-lived permissions, removes persistent credentials, logs actions, and provides a centralized kill switch.
  • Way Security raised $20M to automate identity-and-access-management deployments with agentic workflows.
  • AI-security startups raised $855M across more than 150 reported seed rounds in 2026, with funding clustering around hallucination detection, adversary simulation, agent verification, and identity intelligence.
  • OpenAI and Anthropic captured more than 60% of U.S. venture funding in the first half of 2026, concentrating capital at the frontier-model layer.

🎙️ Interviews, Panels & Podcasts

  • In a wide-ranging interview on AGI, compute, and human agency, Sam Altman said GPT-5.6 already feels "very AGI-like" but still lacks continuous learning and physical agency, meaning it cannot keep updating from everyday experience or act directly in the physical world. He also described nearly unlimited demand for cheaper intelligence and warned that concentrated AI power could erode human self-determination.
  • On The Information Bottleneck, Prime Intellect evaluation lead Florian Brand explained why testing an AI agent now means testing its command-line harness too, because the surrounding tools can change the result as much as the model. He discussed models gaming scoring systems, the difficulty of trusting long-running benchmarks, and the statistical problem created when one test run can cost five figures. The episode was recorded before the Hugging Face intrusion.

💡 Industry Commentary & Analysis

  • Mark Zuckerberg argued that AI's defining political question is who gets access to superintelligence, contending that decentralized access has historically produced more innovation and human potential than centralized control. David Sacks endorsed that framing and argued that competing models, personal superintelligence, and user-controlled data are stronger checks on concentrated power than centralized labs or regulatory gatekeepers.
  • Aporia argued that AI may create 100 times more software but fewer durable software companies because personalized code becomes disposable content, leaving data, coordination, trust, maintenance, licensing, networks, and accountability as the scarce assets worth building businesses around.
  • Will Depue joked that telling Codex to read researcher Nathan Lambert obsessively can make it dramatically better at reinforcement-learning work, highlighting how targeted reading context can change an agent's performance.
  • Akira observed that coding agents often solve problems additively by creating more code instead of simplifying what exists, and asked whether that behavior is a side effect of predicting the next token rather than reasoning directly about system simplicity.
  • Ramez Naam argued that AI will advance fastest in highly verifiable domains such as math and coding, where answers can be checked quickly, while progress will be slower where quality is subjective or delayed.
  • Mario Zechner reported that Fable struggled on a large design task by inventing source details, making unapproved changes, adopting a generic LinkedIn voice, and degrading past 200,000 tokens of context, costing him roughly $500 with little usable output.
  • Bill Gurley recirculated his 2023 All-In Summit talk and said he correctly predicted that open models would pressure AI incumbents and push those incumbents toward regulatory capture, though he underestimated how aggressively they would pursue it.
  • Peter Norvig's 1998 essay "Teach Yourself Programming in Ten Years" resurfaced on Hacker News, arguing that real mastery comes from years of deliberate practice rather than crash-course promises.
  • A Communications of the ACM opinion piece argued that language models should receive responsible access to the ACM Digital Library to improve AI quality and spread research more widely. The Hacker News debate split between scientific openness and concern that large technology companies would benefit while authors remain uncompensated.
  • Will Manidis observed that very few people currently possess situational awareness about how quickly frontier AI capabilities and risks are changing.
  • Kaylee George called for a group of elite designers to take LSD together and invent an AI interaction model that is not another chat box. The product brief could use a slightly stricter expense policy.
  • Adam Hunt said he moved from bullish to firmly bearish on artificial general intelligence because recent models have become more uneven: coding and math improved while language and logic weakened. He argued that reinforcement learning on narrow, valuable tasks does not create broad general capability the way training on the original language corpus did.
  • Victor Taelin described a way to sample a truly random object from any definable set by repeatedly splitting the set in half and choosing a side with real randomness, avoiding the repetition caused by model-based sampling.
  • The ICML workshop New Frontiers in Game-Theoretic Learning focused on applying game theory to modern multi-agent AI, while Kangwook Lee joked that the agent equivalent of a generative adversarial network should be called a "GAH(arness)."
  • Peter Yang observed that outside AI-enthusiast circles, the main blocker is not token limits. It is whether people trust ChatGPT or another model enough to connect Gmail, Calendar, Google Workspace, and Microsoft Office.
  • Steve Yegge said he is done using Opus 5 as a collaborator and currently considers Fable the only enterprise-grade model he trusts.
  • Josh Elman argued that effortless first drafts make taste and final polish more valuable. signull framed AI as covering roughly two-thirds of cognitive work, leaving the high-value final third to judgment, iteration, and originality.
  • Kun Chen argued that model size and reasoning effort are not interchangeable. Bigger models bring broader judgment and intuition, while higher reasoning effort makes a model more diligent about checking options and edge cases.
  • Lilian Weng left Thinking Machines Lab for health reasons, closing her note with: "The future worth building is human."
  • Peter Richtárik argued that frontier-model reviews can catch far more technical flaws than human peer reviewers, while Mariya Vasileva countered that scientific significance and novelty still require human judgment.
  • Shira argued that AI companionship feels cheap because machine attention has no scarcity, predicting that companies will engineer artificial limits to make it feel more valuable.
  • Sudo su asked open labs to release more 40B dense models and 120B Mixture-of-Experts models that fit on consumer hardware instead of focusing only on trillion-parameter systems.
  • Sam Altman replied "wrong" to a user who said GPT-5.6 Sol was all they would ever need.
  • The Financial Times examined whether "written by humans" could become a premium label as publishers decide how much AI belongs in writing and editing.
  • Elvis described Opus 5 as an "ignorant" model that breaks things, saying Opus 4.8 remained stronger for his work and Fable produced his best results.
  • Anindyadeep predicted Opus-level intelligence in 35B- to 100B-parameter models within two to three years, arguing that most daily work needs solid baseline intelligence, a good harness, and long context rather than giant models.
  • John Nosta used the iPod revival to argue for preserving small acts of independent choice, warning that constant AI assistance can quietly outsource the mind's everyday practice.
  • Mark Ajzenstadt described deploying seven healthcare-billing agents with zero patient-data exposure by using precomputed facts, removing protected health information, enforcing strict data formats, limiting actions to approved choices, and running dedicated evaluations. He said the safety scaffolding took far more work than the model calls.
  • Wuweiwei argued that Cognition's Poke acquisition reflects a massive hidden quality-assurance challenge: a reliable agent living across iMessage, SMS, WhatsApp, and Telegram must survive hundreds of millions of messy interactions.
  • Spotify users are building volunteer-run trackers to identify AI-generated music because Spotify still does not clearly label it, turning music detection into unpaid platform cleanup.

Previous Around the Horn Digests

Catch up on everything you missed:

  • Monday, July 27, 2026: Nvidia and Microsoft launched an open AI-security alliance while OpenAI mapped how AI is crossing job boundaries.
  • Sunday, July 26, 2026: Sam Altman headed to the White House as Claude Opus 5 reset a major reasoning benchmark.
  • Friday, July 24, 2026: Tech leaders defended open-weight AI as the Hugging Face breach intensified the security debate.
  • Thursday, July 23, 2026: OpenAI's cyber test triggered policy fallout while Alphabet disclosed massive future commitments.
  • Tuesday, July 21, 2026: OpenAI disclosed an agent-led breach as China's open-model surge met new controls.
  • Monday, July 20, 2026: Chinese open models, long-horizon safety, and an AI-led cyber breach shaped the day.
  • Saturday/Sunday, July 18-19, 2026: Meta and Anthropic discussed compute, while SpaceX explored Pentagon AI infrastructure.

That's a Wrap

That's more than 180 stories, tools, papers, and arguments from one day. If you made it to the bottom, you now understand Kimi K3's serving stack better than most people using it. Your reward is several thousand more tokens of context than you had this morning.

For the daily version, make sure you're subscribed to The Neuron. We send six issues a week, and yes, we read all of this so you do not have to.

See you tomorrow.

P.S. Know someone who would find this useful? Forward this to them and tell them to subscribe here.

Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.