Executive Summary
The last 24 hours saw significant AI model and tooling releases: OpenAI shipped new Codex features including Appshots and token analytics, Zhipu launched GLM-5.1-HighSpeed claiming 400 tokens/s, and Qwen 3.7-Max dropped with strong benchmark performance rivaling Opus 4.7. The open-source agent ecosystem heated up with Teknium's Hermes Agent topping OpenRouter and adding BitWarden key management, while Microsoft reportedly canceled internal Claude Code licenses over cost concerns. On the infrastructure side, llama.cpp gained a new WebGPU backend enabling browser-based inference, and debates intensified around local model deployment as subsidized AI token pricing shows signs of ending.
Key Events
- OpenAI ships Codex updates: New "Appshots" feature, token analytics, and plugin sharing for enterprise users, with Sam Altman and GDB both highlighting the release. → link
- GLM-5.1-HighSpeed released: Zhipu's new model claims 400 tokens/s output speed, with benchmarks showing 300+ TPS in testing — a 10x speed improvement over equivalent models. → link
- Qwen 3.7-Max launches: New Qwen model shows strong visual research and development capabilities, with users noting it beats Opus 4.7 in some tasks. → link
- WebGPU backend lands in llama.cpp/ggml: Enables full-fledged GPU inference directly in the browser, lowering the barrier for on-device AI. → link
- Microsoft cancels Claude Code internally: Token-based billing reportedly made costs untenable even for a major corporation, signaling inference cost pressures. → link
- Hermes Agent claims #1 on OpenRouter: Teknium's general-purpose agent gains traction with BitWarden integration for key management and Honcho-powered memory. → link
- LTX Video model with open weights: 22B parameter video generation model with full trainer released for fine-tuning. → link
- BabyAGI author publishes first academic paper: "The Log is the Agent: Event-Sourced" on arXiv, moving from code citations to formal research. → link
Analysis
Patterns observed: - Agent infrastructure maturing rapidly: Multiple tweets focus on agent payment systems (Revolut cards for agents), key rotation (BitWarden), shared memory layers (gBrain + Honcho), and team coordination — signaling agents are moving from demos to production workloads. - Cost gravity shift: Microsoft canceling Claude Code, warnings about ending subsidized token pricing, and advocacy for local inference (Qwen 3.5 27B on RTX 3090s) all point to an inflection where inference economics can no longer be hidden by venture subsidies. - Browser as AI runtime: WebGPU in llama.cpp and enthusiasm for JS/WASM as a universal runtime indicate the frontier of deployment is shifting toward edge/browser inference, reducing dependence on cloud APIs.
What to watch next: - Whether GLM-5.1-HighSpeed's speed claims hold up under independent testing and whether it changes API pricing dynamics. - The trajectory of OpenAI's Codex agent as it adds enterprise features — this is becoming OpenAI's primary developer-facing product. - Whether "the model is no longer the product" (GDB's quote) accelerates the shift from model companies building their own harnesses to supporting third-party agent ecosystems. - Local model deployment viability as token subsidies recede — Qwen's 27B tier and Hermes' local model support are early signals.
Tweet Feed
AI Model Releases & Announcements
@Ex0byt · 2026-05-21T21:04
Cant wait for ya'll to get your hands on this one.. Latest drop from Qwen:
Qwen3.7-Max-2026-05-20- Looking forward to giving it the full royal PRISM treatment. It's good. → tweet link
@Ex0byt · 2026-05-21T21:28
Qwen3.7-Max beats Opus 4.7 hands down for Visual Research & Development. asked the model to breakdown Derivation of Anchored-SFT Optimization Bounds - visually! → tweet link
@louszbd · 2026-05-22T03:45
RT @zRdianjiao: 🚀 GLM-5.1-HighSpeed is live: 400 tokens/s — a new speed ceiling for flagship-tier LLM APIs. Not a smaller model traded for speed... → tweet link
@TheAhmadOsman · 2026-05-22T18:01
If you saw Qwen 3.5 27B and didn't see that as a chance to ensure the permanent underclass semi-joke never happens to you (by simply purchasing a couple of RTX 3090s) I am sorry to say but you ngmi → tweet link
@victormustar · 2026-05-22T09:17
Made a free Pixal3D demo (Tencent's new image-to-3D model) because I like it a lot 🔥 What's interesting: pixel-aligned generation: every point in the mesh ties back to a specific input pixel... → tweet link
@victormustar · 2026-05-22T13:56
RT @ltx_model: Open weights. A full trainer. 22B parameters you can fine-tune yourself. LTX doesn't just generate video, it lets you… → tweet link
@LinusEkenstam · 2026-05-22T04:29
Insane capabilities of Gemini Omni → tweet link
OpenAI / Codex Updates
@sama · 2026-05-21T20:30
new codex ships today! → tweet link
@gdb · 2026-05-21T19:50
codex app continues to get extremely good, plus features for businesses and enterprise such as token analytics and plugin sharing → tweet link
@gdb · 2026-05-22T12:33
try Appshots in the Codex app: → tweet link
@gdb · 2026-05-22T02:30
codex for using all apps on your computer from your phone → tweet link
@steipete · 2026-05-22T13:05
RT @OpenAIDevs: It's Codex Thursday, and yes, we have updates for you. First up: Appshots, a new way to bring the context of what you're working on into Codex... → tweet link
@TheAhmadOsman · 2026-05-22T04:02
OpenAI, what the fuck is this? Give me back the ability to specify WHICH MODEL I am using + their effort levels EXPLICITALLY. I don't want this router crap you're enforcing on my $200 paid subscription I DID NOT AGREE TO THIS SHIT → tweet link
Hermes Agent & Agent Infrastructure
@sudoingX · 2026-05-22T04:13
hermes agent came like a storm and threw openclaw upside down back to the ocean where it belongs. this isn't a coding agent. this is THE general agent. one agent that does everything. local models. frontier models. autonomous goals. memory. tools. no bloat. no telemetry. no corporate leash. the market already spoke. #1 on openrouter. → tweet link
@Teknium · 2026-05-22T17:45
We got BitWarden now, to make it easy to manage your keys, rotate them quickly, and coordinate access with your team. → tweet link
@Teknium · 2026-05-21T20:27
Anyone tried in Hermes yet? New OS king? → tweet link
@victormustar · 2026-05-22T13:13
what do you use to make your Agents pay online? I was thinking of using a dedicated and limited Revolut card → tweet link
@Teknium · 2026-05-22T17:28
RT @shannholmberg: I've started experimenting with gBrain + Hermes Agent — it's a shared memory layer that sits underneath my Hermes Agent... → tweet link
@sudoingX · 2026-05-21T19:05
anyone here actively running grok build, what are you finding? specifically curious about long running tasks, how the agent holds up over hours, multi file work, and how it compares to whatever you were using before. → tweet link
@thdxr · 2026-05-22T01:29
we setup gangprompt.opencode.ai (threw it behind cloudflare access for SSO) which is running an opencode server on a fast machine with all our repos cloned. now our whole team can pop in there and prompt some stuff and see everything everyone is doing → tweet link
Cost Pressure & Local Inference
@TheAhmadOsman · 2026-05-22T16:39
How it feels to know that the subsidized token industry is coming to an end and people are still not worrying about the right thing (running the models locally on your own compute) No more free DoorDash deliveries / $5 Uber trips for you (after they got you addicted) → tweet link
@sudoingX · 2026-05-22T04:03
which vram tier should i benchmark next? finding the best model for each class. → tweet link
@sudoingX · 2026-05-21T19:54
27b dense is the size serious local llms run on. qwen 3.6 27b owns that tier right now. we need this at 3.7. → tweet link
@jezell · 2026-05-22T05:31
RT @BrianRoemmele: MICROSOFT CANCELED CLAUDE CODE! IT COST TOO MUCH. Major tech companies are confronting the steep reality of AI inference costs... → tweet link
Infrastructure & Developer Tools
@Ex0byt · 2026-05-22T12:51
Glad to see the WebGPU backend making it into ggml and finally getting the moment it deserves! → tweet link
@badlogicgames · 2026-05-22T15:14
i couldn't love this harder. i mean, my old bones are still revolting that JS runtimes + WASM is what we as an industry decided is the new lingua franca. but if it works, it works, and i welcome every new platform where can run inference on easily. → tweet link
@thdxr · 2026-05-22T14:33
had OpenCode run some benchmarks comparing git cli, libgit2, gitoxide, isomorphic-git. even with process spawning overhead, git cli usually won everything. that said i was on linux, overhead on windows probably kills it → tweet link
@levelsio · 2026-05-22T18:25
It's basic common sense. Ports only open for you (via Tailscale) or web traffic (via Cloudflare). Open for 8 billion people or for 2 companies. What would be safer? → tweet link
@levelsio · 2026-05-22T18:17
Yes but block ALL firewall inbound, install Tailscale for SSH and Cloudflare Tunnel on web port for site → tweet link
@levelsio · 2026-05-22T13:20
✨ Added a review system to [site], first just my own reviews. Then next everyone else. But I need a way to avoid fake reviews so I will think long and hard until I open it up for anyone → tweet link
@MengTo · 2026-05-22T16:31
I built a tool to use images 2.0 to generate custom cursors for my screen recordings. → tweet link
@victormustar · 2026-05-22T10:56
I've mostly stopped using Skills and replaced with github gist it's just easier to manage 👀 → tweet link
@jezell · 2026-05-21T22:30
Well, I'm pretty sure I found a gvisor bug today. → tweet link
AI Research & Academic
@jsuarez · 2026-05-22T14:17
This week, I solved a problem in RL involving ludicrous sparsity that I have been thinking about since 2018. Initial sweeps are showing SOTA on one of our most consistently informative test envs. Blog post soon. → tweet link
@badlogicgames · 2026-05-22T13:16
RT @yoheinakajima: babyagi has ~200 citations, but 0 papers... i just published my first paper on arXiv 😆 "The Log is the Agent: Event-Sourced..." → tweet link
@gdb · 2026-05-22T03:51
the model alone is no longer the product → tweet link
AI Quality & Industry Criticism
@badlogicgames · 2026-05-22T17:43
i see this in sectors outside IT. there's a clear race to the quality bottom everywhere now. translator friend now gets to correct AI slop translations. most companies will not even have a translator review AI slop, but throw them at customers verbatim. → tweet link
@badlogicgames · 2026-05-22T15:28
spotted in a WSJ article on vibe slop. → tweet link
@iamdevloper · 2026-05-22T12:58
what I'm gleaning from recent announcements: don't work for a company that's big enough where they have the budget to spend on making your job obsolete with AI → tweet link
@badlogicgames · 2026-05-22T17:50
perfect example to illustrate VLM failure modes. we aren't there yet. → tweet link
@LinusEkenstam · 2026-05-21T20:25
The great cull continues. People have seen nothing yet unfortunately. 10x more impactful than the industrial revolution, 10x faster. Time to have difficult conversations → tweet link
Hardware
@FrameworkPuter · 2026-05-21T23:40
Our first-ever AMD-powered laptop, Framework Laptop 13 with Ryzen 7040 Series is now fully sold out. It had an awesome multi-year run, but the processor is now EOL, so we're unable to make more. Our Ryzen AI 300 Series version is the AMD-based replacement! → tweet link
Developer Practices & Philosophy
@badlogicgames · 2026-05-22T15:20
occassionally write some code by hand, specifically if you learn a new platform/API. let the agent roast you, asking you to explain things to it. pair programming with others on your team, to soak up that institutional knowledge. also easier to make friends that way. → tweet link
@kunchenguid · 2026-05-22T05:50
i'm strongly against model companies focusing too much on harness. if openai didn't build GPT 5.5, no one else can. this is their core competence. if openai didn't build codex cli and app, we have opencode and t3code. the world might be a better place if model companies focus more on their core capability. → tweet link
@TheAhmadOsman · 2026-05-21T20:29
The first LLMs learned from the Open Internet - Blogs, Forums, Repos, Public discourse. Now the BEST coding signal lives inside CLI sessions - Prompts, Diffs, Logs, Corrections, Reactions. The next models get better through YOUR DATA. → tweet link
@gdb · 2026-05-22T06:05
trying to remember what it was like to code before codex → tweet link