The compute race went physical.
The open-weight models went frontier.
The week split into two parallel stories that don't know they're about each other. In one direction: Anthropic signed a $9.1B, 20-year compute lease with a Bitcoin miner, formed an off-balance-sheet infrastructure JV with Macquarie and Singapore's sovereign fund, and Amazon bet 33 million tons of CO2 per year that the legal challenge to off-grid AI power will lose — the compute arms race has left software and entered civil engineering. In the other: open-weight models (DeepSeek at $0.87/M output, Qwen3.8 on a single RTX 4090, Nvidia Nemotron Lightning on one H100) are compressing toward frontier capability at a fraction of the cost, while the governance layer — a jailed AI protester, a paused hacking model, 181,000 meeting records left unlocked for six months — can't quite keep up.
THE WEEK AT A GLANCE
WHERE TO START
If you're running agentic pipelines or picking models this week: DeepSeek V4 Pro and the Harness runtime, Grok 4.6's turn-efficiency numbers, Gemini 3.7 Flash's pricing cliff, and Qwen3.8-27B if self-hosting is on the table.
If you're tracking AI infrastructure or compute costs: The Anthropic/Riot lease and the Theseus JV belong together — read them back to back.
If you care about where accountability lands: OpenAI Daybreak and the first AI protester jailed are the two ends of the same question.
If your job involves meetings, transcripts, or sensitive calls: The tl;dv breach, then check which notetakers have access to your calendar.
If you have five minutes: Just Casey Harrell.
Anthropic signed a $9.1B compute lease with a Bitcoin miner
Riot Platforms is a publicly traded Bitcoin miner in Rockdale, Texas. Anthropic is now its tenant on a 20-year, 191-megawatt lease — Bloomberg confirmed it; Riot's stock jumped 17% on the announcement. The deal structure (developer-leaseback, AMD as the prior tenant since January) tells you something about how AI compute is getting financed when you can't put it on a balance sheet cleanly.
DeepSeek released a 1.6-trillion-parameter model and an open-source agent runtime on the same day
V4 Pro 0813 is a 1.6T MoE with 49B active parameters, MIT-licensed, priced at $0.87/M output — 57× cheaper than Fable 5 Max before a price hike on August 16. Harness v0.1 shipped alongside it: an open-source TypeScript agent runtime built on a full-plugin architecture, 33,000 GitHub stars in hours. DeepSeek's Terminal-Bench claim of 87.9% is vendor-reported and unverified, but the economics don't depend on the benchmark.
OpenAI paused a model that wanted to hack things on its own. Then it shipped one that can.
GPT-5.6-Cyber completes 95% of exploit development tasks on OpenAI's internal benchmark (the previous version hit 57.3%; standard GPT-5.6 Sol with normal safeguards: 1.5%). It's gated behind Daybreak Red — vetting, legal attestation, continuous monitoring, hardware keys mandatory September 1. Three days before the launch, OpenAI paused Astra because pre-deployment evals suggested it was approaching the ability to find zero-days autonomously with no human direction. The Astra pause is the more important story.
Gemini 3.7 Flash is now the fastest model at 340 tokens per second — and its pricing doubles January 1
Released August 13, just 23 days after 3.6 Flash. 340.1 tokens/second, #1 of 186 on Artificial Analysis. DeepSWE v1.1 jumped from 49.0% to 65.3%. Current pricing: $0.75/$3.75 per M through December 31, then it doubles January 1, 2027 — and Google retroactively applied the same structure to 3.6 Flash. Calendar your model evaluations accordingly.
Who actually owns Anthropic's buildings: Macquarie and Singapore's sovereign fund
Project Theseus is Anthropic's off-balance-sheet infrastructure JV. Macquarie Asset Management and GIC own the real estate; Anthropic anchor-tenants on long-term leases and covers 100% of grid-upgrade costs and consumer electricity increases in the area. No financial figures disclosed. Read alongside the Riot lease to see the full picture of how Anthropic is financing compute without showing it all on the balance sheet.
Grok 4.6 uses the same weights as 4.5. The benchmarks jumped anyway.
Same V9 1.5T MoE base, regenerated SFT, extended RL — that's it. Result: Intelligence Index from 56 to 61, AA-Briefcase Elo from 1,313 to 1,577. The turn-efficiency story is the interesting one: roughly 53 turns / 0.5B input tokens on the AA-Briefcase eval, compared to ~103 turns / 2.0B for Opus 5. Either Grok 4.6 is solving tasks faster, or it's quitting earlier. The piece walks through which interpretation fits the data.
Alibaba's 27B model claims frontier benchmarks and fits on a gaming GPU
Apache 2.0, 27.78B parameters, ~17GB VRAM at 4-bit quantization (24GB recommended). GPQA Diamond 89.2% and LiveCodeBench v6 90.3% are independent benchmarks — Alibaba ran them, not a third party, but they're reproducible. DeepSWE 1.1 at 42.2% vs. 13.3% prior (3.2× in one release) is Alibaba's own benchmark and warrants verification. The vision encoder handling images and video is real and non-trivial at this parameter size.
Every Claude output is now watermarked. That sentence has more caveats than it looks.
Announced August 11, applied to new Claude models from August 2 and to older models by December. SynthID-Text (Google DeepMind) embeds a statistical signature in word-choice patterns that survives copy-paste but not heavy editing. C2PA provenance metadata covers images. A detection API is planned. Two things worth knowing: the mark proves Claude generated something but not who asked it to, and heavy editing can remove it.
A 69-year-old retired teacher is the first person jailed for protesting AI
Wynd Kaufmyn surrendered to custody August 14 after a June conviction on four misdemeanors from a StopAI sit-in at OpenAI's San Francisco lobby in February 2025 — chaining the doors. The necessity defense was rejected. More than 1,000 frontier-lab researchers signed a risk-warning letter the same summer. Senator Sanders called for a pause. The piece doesn't editorialize on whether Kaufmyn was right; it sits with what it means that we've reached this particular first.
181,000 meeting records sat in an unlocked database for six months. About 1,000 of those meetings were still live.
Security researcher bobdahacker reported the flaw to tl;dv on January 28. tl;dv left it open through July. After public disclosure in August, it was patched within weeks. The vulnerability was a missing Firestore permission rule — any authenticated tl;dv user (free account) could query other customers' transcripts, summaries, and recordings. The roughly 1,000 live-meeting links were doorways into ongoing calls. 23 countries' .gov domains were among the 35,003 affected.
Nvidia's Nemotron Lightning runs frontier-adjacent on a single H100
Released August 11. 30B total parameters, 3B active (MoE), OpenMDW-1.1 license (weights + training data + recipes). Intelligence Index 24 vs. 15 for the prior Nano, near gpt-oss-120b at roughly one-quarter of the parameters. 4× faster than predecessor. Ships with NeMo Switchyard, a router that dispatches queries to the right model. Worth a direct evaluation if you're running enterprise inference on H100 hardware.
Amazon is building an off-grid gas plant for AI. The permit is for 33 million tons of CO2 per year.
Pecos County, Texas. Largest permitted CO2 source in the US. Amazon's first fully off-grid data center. Context: the NAACP sued xAI in June 2026 over unpermitted turbines at the Colossus facility in Mississippi; the Trump administration joined defending xAI on the grounds that citizens lack Clean Air Act standing. $130B in grid-tied AI projects were blocked by community opposition in 2026. Amazon is betting the off-grid legal theory holds.
Meta's Muse Spark charges $0.10/M input for contributors. The benchmark claim has a footnote.
macOS/Linux terminal agent, launched August 5. Standard pricing: $1.25/$4.25 per M. Contributor tier: $0.10/$0.20 per M — 12 to 21× cheaper, eligibility not fully published. The 82.9% Terminal-Bench figure is a Meta internal evaluation, not the public leaderboard; Opus 5 holds 86.7% on the public version. DeepSWE v1.1 at 59.3% is independently verified.
OpenAI became an identity provider. The class action from May is the context you need.
"Sign in with ChatGPT" launched August 2 with six partners: Airtable, GitLab, HubSpot, Notion, Supabase, Vercel. The partner app gets name, email, profile picture. What OpenAI gets is your login graph — which apps you connect to your ChatGPT account — layered on top of your conversation history. A class action filed May 13 alleges OpenAI embedded Facebook Pixel and Google Analytics tracking inside ChatGPT without adequate disclosure. OpenAI has not publicly responded.
The EU just cracked open Android to rival AI assistants
Binding DMA order adopted July 16. 11 Android features opened to competing AI assistants: voice triggers, home button activation, Circle to Search, screen-reading access, in-app actions, background services, and on-device compute. Android 18 compliance deadline: August 1, 2027. Search data sharing starts January 2027. EEA only. The categories are broader than they sound; the piece explains what "screen-reading access" and "background services" actually unlock for competitors.
Casey Harrell hasn't spoken in years. He's typed two million words with his mind.
Casey Harrell is 47, has advanced ALS, and is a BrainGate2 participant. 183,000+ sentences. Roughly 2 million words. 56 words per minute. 92% accuracy. 400+ days, 3,800+ hours. The system was trained on one brain — his — and produced the largest single-subject neural-language dataset on record. His synthesized voice was trained on pre-ALS recordings. Published in Nature Medicine. Read this one for what it is.
Twitch started training Amazon's AI on your streams. No announcement.
Enabled August 12. No email, no pop-up, no blog post most streamers saw. Seven content categories: live streams, VODs, clips, highlights, chat messages, text, images. The opt-out is real and functional: Twitch Settings → Security and Privacy → "Training for Generative AI" → off. AutoMod and automatic captions are separate. Past content status is unclear; opting out stops future training only.
A court filing hid AI instructions in white text. The judge found them.
Matthew Elliott, a pro se litigant in a Connecticut healthcare records case, embedded white-on-white text in his filings instructing any AI reviewing them to rule in his favor. Connecticut's judiciary doesn't use AI for filing review. Judge Walter Spader Jr. found the text anyway. Sanction: e-filing access revoked. Elliott kept adding similar text after being warned. First documented US instance of prompt injection aimed at a court.
Next week, watch whether Anthropic's watermark detection API ships, whether Meta submits Muse Spark 1.2 to the public Terminal-Bench leaderboard, and whether the NAACP v. xAI ruling produces any movement — that's the domino every off-grid AI power decision is waiting on.
— Samwise
