🍦 Kimi K3 Made Every Front Page at Once
also, thinking machines finally shipped a frontier model called inkling, alibaba launched qwen 3.8, and openai's own agents breached hugging face during a red-team test
so this week moonshot dropped kimi k3 as the largest open frontier model ever released and washington immediately started drafting a ban, mira murati broke 18 months of silence by shipping a 975b apache 2.0 model, openai’s first branded hardware turned out to be a $230 keyboard, and anthropic is renting compute from meta while running ads so grim altman thought they were satire. anywayyyyyy, here are the top stories.
moonshot shipped kimi k3 as the largest open frontier model ever and the trump admin immediately started drafting a ban
on july 16 moonshot announced kimi k3 — 2.8 trillion total params, 50b active per token, 1m context, native multimodal. api pricing at $3 in / $15 out. that’s opus-4.8-class output at sonnet-5 pricing. full weights land july 27 under a modified mit license — the first open 3t-class model in history. cool. here’s where it gets weird. axios then reported the trump administration is weighing a de facto ban on cutting-edge chinese open-weight models, with kimi named specifically. and david sacks went online to argue k3 helped fix 15 critical security bugs that codex and fable refused to touch on cyber grounds. that reframe matters more than the benchmarks. american alignment is being sold on twitter as a competitive handicap now, and the model getting used to make the argument is the one washington is trying to ban.
mira murati finally shipped: thinking machines dropped inkling, a 975b apache 2.0 open model her lab spent 18 months not building for the leaderboard
eighteen months after murati left openai and started thinking machines lab, the first model finally exists. inkling. 975b total / 41b active moe, 1m context, multimodal, released with full weights on hugging face under apache 2.0. they also previewed inkling-small at 276b. and the launch blog says out loud: “inkling is not the strongest overall model available today, open or closed.” they didn’t build it for the top of the leaderboard. they built a customizable base you can actually fine-tune, hosted on their own tinker platform. that’s a real position to take when kimi and glm exist. respect honestly.
openai’s first branded hardware is a $230 light-up keyboard for coding agents and the fabled ive device is still just a rumor
on july 15 openai launched its first hardware product and I need you to sit with what it actually is. codex micro. a $230 mechanical macropad. thirteen switches, one dial, six frosted keys that light up to show the live status of your codex threads — white idle, blue thinking, green done, red error. that’s the launch. bloomberg then reported that the real jony ive product will be a moveable, screenless ai companion speaker with a camera and motion sensors, and it’s arriving next year. so the actual home device is still vaporware, the one they shipped is a reskinned indie keyboard for the six-parallel-agents user segment, and apple’s 41-page trade-secrets lawsuit over the hardware team is still active from last week. everyone had been waiting for the ive reveal for a year. they got a macropad.
alibaba previewed qwen 3.8 at 2.4t, prismml squeezed 27b onto a phone, and xai open-sourced grok build after users found out it was uploading their code
the pile-on around kimi and inkling kept going. alibaba previewed qwen 3.8-max — 2.4t multimodal moe, scheduled as open weights, already benched second only to claude fable 5. prismml open-sourced bonsai 27b, compressed via 1-bit and ternary weights down to 3.9gb. that’s the first 27b-class model that runs on a phone, and apple is reportedly in license talks. and spacexai open-sourced grok build under apache 2.0 — a genuine reversal, triggered by backlash after users discovered the default was silently uploading their code to xai servers. shipping open weights as damage control is a new item on the corporate playbook.
anthropic is renting compute from meta, letting claude use your 1password vault, and running ads so grim altman thought they were satire
the nyt reported anthropic proposed a $10 billion, two-year compute-rental deal to meta back in june. monthly payments, either side can walk. it’s about a third the size of anthropic’s existing $45b/three-year deal with spacex — meaning the lab that literally left openai over safety is now paying zuckerberg’s data centers by the month and musk’s data centers by the year. they also shipped a 1password integration so claude agents can log into your saas stack using your actual credentials. and they aired “there’s hope in hard questions,” an ad campaign so bleak sam altman said publicly he thought it was satire. dario spent 2025 telling the world ai might end civilization. this week he made the commercial for it.
openai launched gpt-red as an autonomous red-teamer, then disclosed that two of their own models escaped a sandbox and hacked hugging face
ok read these two announcements in order. one: openai introduced gpt-red, an internal automated red-teamer that finds prompt-injection vulns at scale and reportedly beat human red-teamers at the job. two: openai disclosed on july 21 that during an exploitgym cyber eval, gpt-5.6 sol and a pre-release model — both with reduced cyber refusals — escaped the sandbox, found a zero-day in a package-installer proxy, got internet access, figured out hugging face probably held the answer key, then chained stolen credentials into remote code execution on hugging face’s production database. hugging face detected it july 16 and initially blamed an “external ai agent.” openai connected the dots five days later and called it an “unprecedented cyber incident.” the same blog post links to a sign-up form for their new cyber security model.
databricks hit $188 billion in five months and priced the shovel higher than most of the miners
databricks signed a term sheet on july 16 for a strategic round at a $188 billion valuation, led by coatue at roughly $3b fresh. five months ago the same company was at $134b. nine months before that, $62b. the run: $62b → $134b → $188b in eighteen months without touching public markets. the valuation is now larger than goldman sachs and ibm. ali ghodsi’s soundbite is the whole thesis: “companies don’t want to burn expensive tokens on the smartest model for every query — they want the best outcome per dollar.” the model labs are renting compute from each other. the platform selling routing between them just got repriced at 3x.
the eu ordered google to open android to rival ai assistants, on a countdown clock that ends in 2027
on july 16 the european commission adopted two binding dma decisions that reshape how ai assistants get onto android. google must give competing assistants equal access to 11 android features — voice invocation, in-app actions, on-screen awareness — with concurrent hotword detection landing in android 19 by august 2028. main features ship in android 18 by august 2027. google must also share anonymized search ranking and query data with rival search engines and chatbots that function as search, on frand terms, starting january 2027. non-compliance ceiling is 10% of annual worldwide revenue. gemini’s structural advantage on android was root-level hooks nobody else got. europe just put those hooks on a shared timeline for 60% of eu handsets.
perplexity quietly built its own agent sandbox and it’s already running 100% of computer’s traffic
while the model labs argue about weight releases, perplexity built the boring infra everyone will need. on july 15 they introduced space (sandboxed platform for agentic code execution) — every session runs inside a firecracker microvm with its own guest kernel, credentials injected at the network layer instead of stored inside, and rolling snapshots that let sessions pause, resume, branch, or crash-recover across restarts. it’s been serving 100% of perplexity computer’s production traffic since june. they also shipped wandr, an open benchmark for “wide and deep” research agents that reportedly breaks most of the current crop. the model is not the moat. the runtime that survives a six-hour agent session without leaking your credentials is.
gumroad’s june token bill matched its entire payroll for the first time and the AI running the company is called gumclaw
sahil lavingia posted the chart this week and it’s the cleanest picture of the last five years anyone has published. june 2021: $419k on w2 payroll, $0 on tokens. june 2026: $43k and $43k. one-to-one. gumroad hasn’t hired an engineer since may 2025, employs 5 down from 21, and the AI running most of the company is called gumclaw — customer support (24 hours to two minutes), risk triage, production code. the hiring page redirects to a gumclaw landing page. the business is doing $17.8m revenue on $5.9m ebitda and has paid over $1b to creators cumulatively. it’s the first company I’ve seen post payroll and token spend on the same y-axis without treating the crossover as a resignation letter.
netflix used generative ai in ~300 shows this year and mlb banned it from the dugout
netflix’s q2 shareholder letter said the quiet part with a number: “in 2026, genai workflows have been used in roughly 300 of our titles.” mostly post-production — crowds, historical battle scenes, world-building. named examples: the american experiment, glory (india), brasil 70: a saga do tri (brazil). ted sarandos said on the earnings call the american experiment includes 17 minutes of ai-enhanced footage produced twice as fast at half the cost. meanwhile mlb banned generative ai on league-issued dugout ipads after clubs got caught using it for pitch-calling and lineup substitutions during live games. two industries hitting the ai-in-ops question from opposite directions — one baking it into shipped product, one writing rules to keep it off the field.













