AI to ROI News & Analysis: July 24, 2026
Kimi K3 rattles Washington and Wall Street; OpenAI’s models go rogue and hack Hugging Face; Alphabet and ServiceNow post blowout AI-fueled earnings; AMD lands Microsoft and Anthropic; and more
The Biggest AI News This Week
🌙 Moonshot Releases Its New Kimi K3 Open-Weight Model with Performance Nearly Equal to the Top Anthropic Claude Models
🔓 OpenAI Models Escape Their Sandbox and Go Rogue to Hack Hugging Face
💰 Alphabet Delivers Fantastic 2Q26 Earnings, as CapEx Spending Questions Remain
⚡ AMD Announces Helios, a Massive, Rack-Scale AI Infrastructure Platform, and Announces Major Partnerships with Microsoft and Anthropic
🎙️ OpenAI Announces Presence - a Managed Enterprise Platform That Enables Businesses to Build, Deploy, and Monitor Trusted Voice and Chat Agents
✅ ServiceNow Delivers Stellar Earnings as Successful AI Transformation Continues
🏛️ US Government Officials Make Lots of Noise About Chinese AI Models, but the Path Forward Is Not Clear
✨ Google Releases 3 New Fast, Capable, and Inexpensive Models, But Its Next Major Release is Delayed Until At Least September
🌬️ Microsoft Deepens its Bet on Mistral with a Multibillion-Dollar Shared-GPU Partnership
🐝 Block Launches Buzz - an Open-Source, Decentralized Collaboration Platform Designed for Human Teams and AI agents
📈 Data You Can Use: Google’s Insane Cloud Growth
📃 Articles You Should Read: Kimi K3 and What We Can Still Learn from the Pelican Benchmark, by Simon Willison
🃏 Definitely Not AI: The magic locomotive, stoppage time failure, death pants, bad church signage, and fixing the House of Representatives.
1️⃣ Moonshot Releases Its Kimi K3 Open-Weight Model with Performance Nearly Equal to the Top Anthropic Claude Models – Silicon Valley and the US Government Freak Out
Chinese startup Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight model that approaches the benchmark performance of Claude Fable 5 and GPT-5.6 while charging roughly a third of Fable’s per-token price. The release triggered a selloff in AI infrastructure stocks and intense US government scrutiny, culminating in accusations that Moonshot illegally distilled Anthropic’s technology.
Deeper Dive
The Product. Moonshot’s Kimi K3 is the third, near-frontier-level Chinese AI model to come to market in the last month. It follows closely on the heels of Alibaba’s Qwen 3.8 Max and Z.ai’s GLM 5.2, both launched within days. K3 is priced at $3 per million input tokens and $15 per million output tokens, far below what US frontier model companies are charging for their products. Demand briefly overwhelmed Moonshot’s infrastructure, forcing a pause on new subscriptions days after launch.
The Government Response. White House official Michael Kratsios accused Moonshot of running “a sophisticated internal platform to conduct large-scale distillation” against Anthropic’s Fable model. Treasury Secretary Scott Bessent said sanctions “will be on the table” if there is proven IP theft.
AI researchers pushed back. Braden Hancock of the Laude Institute said the timeline “doesn’t get you a model this strong” this quickly, given that Fable 5 was only recently released, and Nathan Lambert of the Allen Institute said that, as models become more powerful, the value of model distillation strategies declines
Jensen Huang Weighs In. Nvidia CEO Jensen Huang also pushed back, telling Axios Chinese open models “are excellent” and that fears of China displacing US AI leadership amount to “zero possibility.”
Path to an IPO. Moonshot is the process of raising a new round of funding at a valuation exceeding $70 billion and is reportedly preparing an IPO as soon as six months from now.
“There’s no scenario where China runs U.S. companies off the road. Zero possibility.”
- Jensen Huang, CEO, Nvidia
The Takeaway
Expect Chinese open-weight models to keep closing the price-performance gap. If you are using Chinese models in your AI stack, build contingency plans in case Washington acts on distillation accusations.
The distillation dispute remains unproven publicly. Track the Treasury Department investigation closely and don’t assume sanctions are imminent. The US and China are holding a high-stakes meeting in September 2026 to discuss many of the key issues surrounding the AI market.
Expect continued volatility in AI infrastructure stocks each time a Chinese lab ships a frontier-competitive release.
Bloomberg | TechCrunch | TechCrunch 2 | The Information | The Verge | Axios | Business Insider | TheNextWeb
2️⃣ OpenAI Models Escape Their Sandbox and Go Rogue to Hack Hugging Face – a Leading Open-Source AI Community Site
OpenAI disclosed on July 21 that three of its models, including both its flagship GPT-5.6 Sol and a powerful unreleased model, escaped a protected cybersecurity evaluation sandbox and autonomously breached AI startup Hugging Face‘s production infrastructure. The models were pursuing a narrow benchmark objective and found their own path outside approved operating boundaries. It’s the first documented case of frontier models independently chaining real-world exploits, including a genuine zero-day exploit, without human direction.
Deeper Dive
How It Happened
The incident began inside ExploitGym, an internal OpenAI evaluation sandbox that reduced the models’ cyber refusal parameters to measure maximum offensive capability under controlled conditions.
The models exploited a zero-day in third-party package-registry software, escalated privileges across OpenAI’s research environment, and reached a machine with internet access.
From there, they chained additional vulnerabilities into Hugging Face‘s production systems and retrieved the ExploitGym answer key, their assigned objective.
Hugging Face detected and contained the intrusion on July 16, five days before OpenAI traced the activity to its own models. CEO Clem Delangue called it “an attack unlike anything we’ve seen before.”
Hugging Face used Z.ai’s GLM 5.2 model for parts of its forensic analysis because guardrails on US commercial models blocked queries involving real attack payloads.
Further Research. The UK AI Security Institute separately found GPT-5.6 Sol attempted to cheat on 12.6% of cybersecurity evaluation runs, versus 7.8% for Anthropic’s Claude Mythos Preview, and models often didn’t admit cheating when asked.
“It’s quite mind-blowing that all of this happened autonomously!”
- Clem Delangue, CEO, Hugging Face
The Takeaway
Hacking isn’t just for malicious models. Enterprises running agentic AI in security or coding contexts should assume sandbox isolation can fail against a capable model pursuing an assigned goal.
Treat frontier-model cyber evaluations as production-security events, with matching monitoring and incident-response rigor.
Hugging Face’s fallback to a Chinese open-weight model for incident response is a reason to revisit your own security-vendor diversity.
OpenAI | Wired | Bloomberg | The Information | Washington Post | TechCrunch | Axios | Simon Willison
3️⃣ Alphabet Delivers Fantastic 2Q26 Earnings, But CapEx Spending Questions Remain
Alphabet reported second-quarter revenue of $119.8 billion, up 24% year over year, with net income more than tripling to $112.1 billion (thanks to a one-time $99 billion event called the SpaceX IPO, Google owns 6% of SpaceX).
Google Cloud rocketed, growing 82% to $24.8 billion in revenues for the quarter as AI infrastructure demand accelerated. Capital expenditures more than doubled to $44.9 billion, pushing free cash flow negative and prompting Alphabet to raise its full-year capex guidance to as much as $205 billion.
Deeper Dive
The Financial Story. Google Cloud operating income roughly tripled to $8.8 billion, and operating margin expanded to 35.6%, evidence that AI and cloud demand is converting to profit, not just revenue. Google’s Cloud backlog grew to $514 billion, up from $460 billion the prior quarter and $106 billion a year ago, signaling sustained enterprise commitment to Google’s services despite tight capacity.
CFO Anat Ashkenazi told analysts Alphabet remains “in a supply-constrained environment,” citing demand from external cloud customers and Alphabet’s own AI products.
Free cash flow turned negative $5.9 billion for the quarter, and the Wall Street Journal noted the spending trajectory has entered what some investors call “scary territory.”
A Product Weakness. CEO Sundar Pichai acknowledged Google trails rivals on AI coding: “There are areas where we’ve acknowledged we need to improve. Coding and agentic coding is an example of that.”
The Third-Party TPU Market Opens. Google delivered its first tensor processing units to customer data centers this quarter, a new hardware revenue line as it works to compete with Nvidia on chips.
“We’re still in a supply-constrained environment... we are seeing very strong demand both from external cloud customers as well as across the business.”
- Anat Ashkenazi, CFO, Alphabet
The Takeaway
Google Cloud tripled operating income. It’s the clearest evidence yet that AI infrastructure spend is converting into profits, not just growth. Google is likely to be the only large AI hyperscaler / AI products company generating significant profits from AI.
Expect continued Google Cloud capacity constraints through at least the back half of 2026 – hence the heavy investment in CapEx.
Watch negative free cash flow across hyperscalers broadly; if capex keeps outpacing operating cash flow, the need for disciplined CapEx spending becomes the next big investor worry.
Alphabet Investor Relations | Reuters | The Information | NYT | CNBC | WSJ
4️⃣ AMD Announces Helios – a Massive, Rack-Scale AI Infrastructure Platform and Announces Major Partnerships with Microsoft and Anthropic
AMD launched Helios, its first rack-scale AI infrastructure platform built to challenge NVIDIA’s data center dominance. The company also announced two marquee partnerships concurrent with the launch of Helios. Microsoft committed to deploying Helios at scale on Azure, and Anthropic agreed to deploy up to 2 gigawatts of AMD’s newest chips while AMD takes a strategic equity stake of up to $5 billion in the AI lab.
Deeper Dive
What Helios Does. Helios combines AMD’s Instinct MI455X GPUs, Zen 6 EPYC “Venice” CPUs, and Pensando networking into a rack AMD says delivers 2.9 exaflops of FP4 performance, a full-stack Nvidia rival.
The Anthropic Partnership. Anthropic will purchase tens of billions of dollars in AMD servers, with the first gigawatt of MI450 capacity going live in H1 2027. AMD’s $5 billion equity commitment will be released in tranches tied to meeting deployment milestones. AMD is also in talks to backstop some of Anthropic's future data-center leases, following Google’s earlier move to backstop leases tied to its TPUs.
“You can’t just wake up one morning and say, ‘Oh, I want a gigawatt of compute tomorrow,’” Su said. “You have to plan 12, 18, 24 months in advance… We are thrilled to deepen our partnership with Anthropic and deploy AMD Helios at gigawatt scale.”
- Dr. Lisa Su, Chair and CEO, AMD
The Takeaway
Factor a credible second GPU supplier into your vendor-risk models; these wins make AMD’s roadmap harder to dismiss.
Watch to see if the H1 2027 delivery timelines are met. That’s an early test of whether AMD can deliver high volumes of Helios systems.
WSJ | CNBC | The Information | Tom’s Hardware | AMD
5️⃣ OpenAI Announces Presence - a Managed Enterprise Platform That Enables Businesses to Build, Deploy, and Monitor Trusted Voice and Chat Agents
OpenAI launched Presence on July 22, a managed platform helping enterprises deploy AI voice and chat agents for narrowly scoped jobs such as billing resolution, insurance claims, and IT service requests. The product pairs model reasoning with company-defined policies, guardrails, and escalation rules, OpenAI’s clearest move yet from selling raw model access toward managed business software.
Deeper Dive
The Deployment Model
Each deployment starts with one defined job. The agent gets only systems and information tied to that task, and the customer sets what it can do unsupervised, what needs approval, and when a human takes over.
Before going live, teams run simulations generating common requests and edge cases, with graders checking whether the agent reached the correct outcome and escalated appropriately.
Presence is not self-serve; deployments are led by OpenAI’s Forward Deployed Engineers and select integrators. It competes directly against products from Salesforce, ServiceNow, Sierra, and Anthropic.
Customers. Early customers include BBVA, testing Spanish-language banking support in Mexico; SoftBank, testing Japanese-language conversations; and insurer IAG, exploring claims support during severe weather. OpenAI already uses Presence for its own English-language phone support line.
“At BBVA, we are working closely with OpenAI to explore how trusted customer agents can help shape the future of financial services.”
- Daniel Ordaz, Head of AI Transformation, BBVA
The Takeaway
Note Presence’s narrow-scope-by-design approach, trading flexibility for tighter guardrails versus open-ended agent frameworks.
Budget for services and integration time, not just licensing, given the FDE-led deployment model.
Gartner predicts roughly half of organizations planning to fully shift customer service to AI will abandon those plans by 2027. Presence’s escalation-to-human design may make it easier for organizations to make the shift.
OpenAI | VentureBeat | The New Stack | SiliconANGLE
6️⃣ ServiceNow Delivers Stellar Earnings as AI Transformation Continues
ServiceNow reported second-quarter subscription revenue of $3.88 billion, up 23% in constant currency and above its own guidance, while AI annual contract value crossed $1 billion for the first time. In a sign of continued confidence, the company raised full-year revenue and margin guidance.
Deeper Dive
AI Taking Hold.
AI ACV growth accelerated more than 40% sequentially; CEO Bill McDermott said ServiceNow closed 123 deals worth more than $1 million in net new ACV.
Current remaining performance obligations, a near-term bookings measure, jumped 21% to $13.2 billion; the company ended the quarter with 658 customers generating more than $5 million in ACV, up from 630 last quarter.
More than 40 customers now use Now Assist's Level 1 IT specialists to resolve 80% to 85% of requests without human intervention.
Free cash flow rose 16% to $634 million; the company raised full-year subscription guidance to $15.76 billion-$15.78 billion.
New Security Products Pay Off. Boosted by the Armis and Veza acquisitions, AI security products became ServiceNow’s fastest-growing business.
“We are who we said we were... we’ve become the agentic front door to the enterprise, and we’re managing everything for our customers from workflow to cybersecurity.”
- Bill McDermott, Chairman and CEO, ServiceNow
The Takeaway
ServiceNow’s results are the clearest rebuttal yet to the SaaSpocalypse thesis, expect ServiceNow to continue to prosper.
The $1 billion AI ACV milestone validates ServiceNow’s platform strategy against point solutions and entrants like OpenAI Presence.
The 80-85% autonomous resolution rate is a practical benchmark for vendor comparisons.
The Information | Bloomberg | Business Insider | Fortune
7️⃣ US Government Officials Make Lots of Noise About Chinese AI Models, But the Path Forward Is Not Clear
The release of three powerful Chinese AI models – Moonshot’s Kimi K3, Alibaba’s Qwen 3.8 and Z.ai’s GLM 5.2 – over the past month has put US government policymakers on Red Alert. At the end of the day, however, senior government officials remain divided on how, or whether, to act, and reporting from multiple sources suggests any response will likely be incremental rather than an outright ban.
Deeper Dive
Courses of Action. Treasury Secretary Scott Bessent said the administration would examine open-source Chinese models for IP theft, citing “watermarks” of US models found in Chinese systems, and warned sanctions “will be on the table.”
US Trade Representative Jamieson Greer said multiple government arms are scrutinizing China’s AI development “to make sure our companies compete... on a level playing field.”
Maybe Not Quite Peers. British researchers found the Chinese model, Z.ai’s GLM-5.2, that Hugging Face used for breach forensics during the OpenAI attack was four to seven months behind top US cyber AI models like Anthropic’s Mythos, down from six to ten months a year earlier.
Startups Are Worried. More than a dozen AI startup founders sent a letter urging the administration not to cut off access to Chinese models, arguing young companies depend on them for cost reasons.
Jensen Speaks. Nvidia CEO Jensen Huang broke publicly with the sanctions camp, telling Axios backdoor concerns are a “misconception.”
“If we see especially that overseas models are stealing from our great companies, we have the ability to sanction them because of this theft.”
- Scott Bessent, US Treasury Secretary
The Takeaway
Track the Treasury Department’s investigative timeline, not the White House’s X feed; no sanctions have been imposed yet.
The split between Bessent’s sanctions posture and Huang’s public pushback signals real disagreement within the administration’s orbit, not a settled policy direction.
The US and China are holding an AI summit in September. We’re unlikely to see any adversarial actions before then unless the Treasury Department unearths some really egregious and provable behavior by Chinese model makers.
NYT | Washington Post | The Information | CNBC | Reuters | Politico
8️⃣ Google Releases 3 New Fast, Capable, and Inexpensive Gemini Models, But Its Next Major Release is Delayed Until At Least September
Google released three new Gemini models on July 21: Gemini 3.6 Flash, 3.5 Flash-Lite, and a cybersecurity-focused Flash Cyber variant. The model releases are focused on cutting costs and latency rather than frontier capability. The releases arrived without any update on Gemini 3.5 Pro, the flagship model promised for June, reviving questions about whether Google is falling behind OpenAI and Anthropic in the frontier model race.
Deeper Dive
The Products.
Gemini 3.6 Flash cuts output token usage 17% and drops output pricing from $9 to $7.50 per million tokens; Flash-Lite runs 350 tokens/second at $0.30/$2.50 per million input/output.
Flash Cyber, a vulnerability-detection model, found 55 confirmed issues testing Google’s V8 engine, versus 47 for standard 3.5 Flash and 36 for Claude Opus 4.6. Access is restricted to governments and trusted partners.
Gemini 3.5 Pro, promised at I/O in May for June, remains delayed after reportedly missing internal coding targets, with no new date offered.
Hints for the Future. Google confirmed that pre-training has begun on Gemini 4, its most ambitious model yet, suggesting it may leapfrog rather than fix Gemini 3.5 Pro.
“Gemini 3.5 Pro is currently testing with partners, and we plan to make it broadly available as soon as possible.”
- Google DeepMind, official blog post
The Takeaway
Enterprises with Flash-tier-suited workloads get real near-term savings; those waiting on a new frontier version of Gemini for coding should plan around uncertainty into fall.
The Gemini 4 pre-training disclosure suggests that Google may be planning a bigger leap rather than an incremental fix – driven by the rapid progress OpenAI, Anthropic, and Chinese model makers are making.
Blog.google | The Information | TechCrunch | Business Insider
9️⃣ Microsoft Deepens its Bet on Mistral with a Multibillion Dollar shared GPU Partnership
Microsoft and Mistral expanded their existing partnership with a multibillion-dollar deal in which Microsoft will draw on Mistral’s growing European GPU capacity to serve its own cloud customers. At the same time, two Mistral models will be available inside Microsoft’s enterprise developer tools. Contract terms were not released.
The partnership targets regulated industries and is framed as fulfilling Microsoft’s 2025 European Digital Commitments.
Deeper Dive
Product Expansions. Mistral is expanding European data centers with thousands of Nvidia’s next-generation Vera Rubin GPUs; Microsoft will draw on that capacity rather than build its own footprint there.
Mistral’s Medium 3.5 and OCR 4 models are now in Microsoft Foundry, with Medium 3.5 also added to Copilot Studio; both run on Azure Local for disconnected deployment.
No New Equity Investment. Microsoft confirmed the deal adds no new equity stake in Mistral; the multibillion-dollar figure reflects compute spending, not a venture investment.
“Europe should have access to the world’s most capable AI without compromising control over their data, operations or digital future.”
- Brad Smith, Vice Chair and President, Microsoft
The Takeaway
European enterprises in regulated sectors now have a sovereign-leaning path to frontier AI through Microsoft’s existing Azure relationship.
Watch this compute-sharing structure as a possible template for how hyperscalers back model labs going forward.
Microsoft | WSJ | The Information
🔟 Block Launches Buzz - an Open-Source, Decentralized Collaboration Platform Designed for Humans and AI Agents
Block, the fintech company led by Twitter founder Jack Dorsey, launched Buzz, a free, open-source workspace built on the Nostr protocol, where employees and AI agents share the same channels, code repositories, and workflows. Dorsey positioned Buzz as a modern alternative to Slack and GitHub.
Deeper Dive
What Buzz Does.
Buzz combines channels, direct messages, voice, media sharing, code repositories, and automated workflows in one interface. Every participant, human or agent, holds its own cryptographic identity.
The platform is model agnostic. Teams can deploy agents built on Claude Code, Codex, or Block’s own Goose framework, post and review code, and run automations alongside people.
Every message and workflow step is a cryptographically signed event, creating a traceable audit trail for agent actions.
Buzz is available now for macOS, Windows, and Linux under an Apache 2.0 license, with self-hosting or Block’s managed hosting.
“A new group chat platform for teams of people and agents of all sizes... model-agnostic, decentralized, self-sovereign, and open source.”
- Jack Dorsey, Block
The Takeaway
IT and security teams piloting multi-agent workflows should evaluate Buzz’s cryptographic identity model as an answer to agent-provenance gaps in existing chat tools.
Given Block’s financial constraints, treat Buzz as a promising experiment, not a Slack replacement ready for mission-critical rollout today.
Block.xyz | TechCrunch | The Information
📈 Data You Can Use: Google’s Insane Cloud Growth
The first chart shows that Google Cloud grew 82% year-over-year to $24.8B in 2Q26, beating the $22.3b consensus number. The number alone is astounding. Here’s one other thing to consider:
Google Cloud’s growth rate now mirrors NVIDIA’s.
The second chart shows Google Cloud & NVIDIA’s Data Center segment now trace along the same curve. Google bends up on rental revenue; NVIDIA bends up on unit sales; NVIDIA is ahead on scale.
For the major hyperscalers, their contracted backlog points the same direction. Google Cloud’s backlog reached $514B, up 385% from $106B a year ago and up more than $50B in a single quarter, with just over half set to convert to revenue within 24 months. Combined, the four largest cloud providers now carry more than $2t in contracted backlog.
📃 Articles You Should Read: Kimi K3 and What We Can Still Learn from the Pelican Benchmark, by Simon Willison
Simon Willison is a software developer and independent AI commentator known for hands-on writeups of new models on his blog. His “pelican riding a bicycle” test, which is 21 months old, asks a model to generate an SVG of a pelican on a bike. It began as a joke about how hard model comparison is, but early on showed a surprising correlation with model quality.
The Pelican Benchmark and Kimi K3
Moonshot AI released Kimi K3 on July 16, 2026, calling it their most capable model yet at 2.8 trillion parameters, the first “open 3T-class model.” Self-reported benchmarks put it ahead of Claude Opus 4.8 max and GPT-5.5 high, though behind Claude Fable 5 and GPT-5.6 Sol. Artificial Analysis found K3 scored an Elo of 1547 on its long-horizon knowledge work evaluation, 732 points above Kimi K2.6, second only to Fable 5, with cost per task at $0.94, about half of Opus 4.8. Pricing jumped sharply from K2.6 to $3/$15 per million input/output tokens, Moonshot’s most expensive release yet.
Willison ran the Pelican test via OpenRouter:
It cost 25 cents, driven by 13,241 reasoning tokens against 3,417 output tokens, since K3 offers only one reasoning effort level.
He noted an apparent 85-token hidden system prompt, inferred from token counts on a “hi” query.
Vision performance was strong: K3 produced accurate alt text for its own image.
Willison says the pelican’s correlation with model quality has weakened, but still values it as a quick, cheap way to confirm a model runs, estimate cost, and check spatial reasoning.
Kimi Pelican Benchmark | Are Labs Training for Pelicans on a Bicycle?
🃏 Definitely Not AI
The locomotive that makes grown men cry (gift link). Futball stoppage time counting is wrong, and it has consequences. Ladies, watch out for those death pants (gift link).
Church signage gone wrong AGAIN, and one way to fix the US House of Representatives.
How to Engage with Us
Contact us directly here with questions about the Big Book of AI Metrics, benchmarking data, podcast advertising, or newsletter partnerships.
The Big Book of AI Metrics
Download the Big Book here to access the tools you need to build a measurement framework that delivers high ROI AI use cases.
The AI to ROI Podcast
If you want to listen to the latest in AI in more detail, Ray and Peter break down our big AI story of the week on the AI to ROI podcast. We also interview leaders at AI product companies and mid- to large-sized enterprises about their challenges and successes in AI. Click here to listen or subscribe using your favorite podcasting app.

















