Daily Intelligence Briefing — Tech / AI / IT Monitor
Date: 2026-05-23 | Reporting Period: Last 24 Hours
Executive Summary
The dominant theme is the accelerating real-world deployment of AI coding agents, with multiple high-profile developers reporting production-grade workflows centered on Codex, Claude Code, and Cursor CLI. Meanwhile, several experienced engineers are publicly noting that AI model capability gains have plateaued since late 2025, suggesting an inflection point in the S-curve. The open-source AI movement gained rhetorical momentum, with concerted pushback against OpenAI's apparent push into user financial data and lock-in via proprietary agent harnesses. On the tooling front, Hermes Agent (from Nous Research) appears to be preparing for a significant launch, while KiloClaw emerges as a hosted alternative for the fast-growing OpenClaw ecosystem.
Key Events
-
@badlogicgames (libGDX creator) published a detailed assessment that AI coding models have hit an S-curve plateau, noting no meaningful step-change since the 3.x→4.x transition in late 2025 and that benchmark improvements don't translate to real-world coding gains. → link
-
@steipete built a cloud-hosted Codex runner on Cloudflare Firecracker boxes with Ghostty via WebAssembly, and separately released an autotriage skill for Codex that uses VM + computer vision to autonomously verify issue/PR fixes. → link → link
-
@gdb confirmed GPT-5.5 is "a very good model" and demonstrated Codex building and debugging an iPhone simulator end-to-end. → link → link
-
@kunchenguid posted a widely-discussed thread on model vs. harness dynamics, arguing that model companies locking users into 1st-party harnesses harms the ecosystem, while praising OpenAI for keeping Codex CLI open-source and allowing 3rd-party harness integration. → link
-
@TheAhmadOsman flagged that a 500B NVFP4 model is coming, warned users about OpenAI requesting financial account data, and shared an extensive free LLM fundamentals resource. → link → link
-
@sudoingX reported that Cursor CLI with Composer is dramatically faster than Claude Code and is the current "killer combo" for shipping code. → link
-
@Teknium announced Anthropic has granted Nous Research security privileges with Claude, and Hermes Agent gained 256K context length, signaling an imminent major release. → link → link
-
@levelsio reported he hasn't written code manually in 6 months, ships all changes live to production via Claude Code, and noted big tech is now restructuring like solo AI builders (firing HR/middle management, merging roles). → link → link → link
-
@RealGeneKim endorsed KiloClaw as a hosted OpenClaw onboarding service, highlighting its value vs. the multi-hour ordeal of self-hosting. → link
-
@TheAhmadOsman called out MiniMax-M2.7 running on 4× M5 Max MacBooks as "performative local AI grifting," with only ~3K max context and 45 tok/s. → link
Analysis
S-Curve Plateau Signal: Multiple independent voices (@badlogicgames, @dhh via RT, @kunchenguid) are coalescing around the view that frontier model capability has flattened since approximately Q4 2025. This is significant because it suggests the next leap may require architectural or data paradigm shifts rather than incremental scaling. If the plateau holds, competitive advantage shifts from model quality to harness/agent design and ecosystem lock-in.
Agent Harness Wars: The @kunchenguid thread crystallized a key tension: model companies are heavily marketing their proprietary harnesses (Codex, Claude Code, etc.) as revenue generators rather than publishing them as specs. OpenAI's semi-open approach with Codex CLI is being positioned as the pro-ecosystem alternative to Anthropic's more locked-down strategy. Expect this to intensify as margins on raw model access compress.
Open-Source AI as Ideological Movement: @TheAhmadOsman and @Teknium are visibly coordinating a narrative push framing open-source AI as a civilizational infrastructure issue (agency, sovereignty) rather than just a licensing question. Combined with the Hermes Agent momentum and Bitwarden security integration, Nous Research appears to be building a full-stack open alternative.
Developer Tooling Rapid Cycling: Cursor CLI → Codex → cmux workflows are evolving weekly. @steipete's Firecracker-based cloud Codex and @sudoingX's Cursor CLI praise indicate the market hasn't converged on a single agent interface. The speed of iteration suggests the winning tool will be whichever best supports worktree-based multi-session workflows (per @thdxr).
What to Watch: (1) Whether GPT-5.5's reported quality gains translate to sustained developer preference or are benchmark artifacts; (2) Nous Research Hermes Agent full launch timeline; (3) Any OpenAI response to the financial-data-access backlash; (4) Signs of model companies adjusting harness-lock-in strategy after community pushback; (5) The 500B NVFP4 model announcement and required hardware implications.
Tweet Feed
AI Model Capabilities & Plateau Debate
@badlogicgames · 2026-05-23T18:32
this has been my experience as well. there definitely were improvements, specifically wrt shell based computer use, but also regressions, especially in the last 3 version bumps of flicker and gerpertee corp. the last step change was 3.x to 4.x in flicker land, probably mostly due to them getting all coding sessions from april to october 2025 via CC. similar timeline with GPT and Codex. at least in my line of work, no big jumps after that. the benchmark increases mean literally nothing in the real world. i suppose we have a data problem now. only so much you can RL into those damn things. and with ralph loops/swarms/agents reviewing agents/whatever, you get less and less human signal to improve RL, would be my uneducated guess. also very hard to capture design/system thinking in RL would be my guess. all that said: if we are at the top of the S curve now, then i'll take what we got. plenty useful, even if it won't replace me fully nor partially anytime soon. → tweet link
@gdb · 2026-05-23T04:51
GPT-5.5 is a very good model → tweet link
@badlogicgames · 2026-05-23T09:33
me: do it / gpt: totally did it / me: dude / gpt: totally did it now / me: wtf / gpt: i so did it, you won't believe how hard i did it — gpt 5.5, thinking off. → tweet link
@TrungTPhan · 2026-05-23T16:57
Claude Opus vs. Claude Mythos → tweet link
@TheAhmadOsman · 2026-05-23T17:49
We're getting a 500B NVFP4 model soon guys. Get your RTX PRO 6000s ready → tweet link
AI Coding Agents & Developer Tools
@steipete · 2026-05-23T18:07
Still limited by compute, so I built a thing that runs codex in the cloud, powered by @Cloudflare firecracker boxes (and since that's not beefy enough for larger projects, tests are run via crabbox). Uses Ghostty ofc, via WebAssembly. Codex replicated itself, basically. → tweet link
@steipete · 2026-05-23T17:36
I built an autotriage skill for codex that has a set of guidelines + reads VISION.md from my repos, so issues/prs that have a clear way of - fit vision of the project - being inferrable in code with high confidence - clear fix - can be live tested Are now worked on autonomously. Codex can use a VM + computer vision (via new parallels backend) to verify fixes, so it can work without interrupting me. → tweet link
@gdb · 2026-05-23T17:05
Codex for building and debugging an iPhone simulator end to end: → tweet link
@sudoingX · 2026-05-23T11:33
cursor cli is so fucking fast it's unreal. if you jumped from claude code the difference is not even close. the speed alone changes how you think about prompting. i'm shipping faster than i ever have. → tweet link
@sudoingX · 2026-05-23T08:02
cursor cli with composer is the killer combo rn. → tweet link
@steipete · 2026-05-23T07:51
I'm late to the party, but cmux is great. current split: codex mac app: knowledge work, learning, reading; cmux + codex cli: coding → tweet link
@sudoingX · 2026-05-23T12:42
finally getting supergrok tonight to try grok build for the first time. the benchmarks say it's behind claude code and codex but i want to feel it myself. benchmarks don't tell you how a tool thinks. grinding it tonight. → tweet link
@thdxr · 2026-05-22T20:50
what pushed us to finally implement it is heavier use of worktrees. now that i have sessions going across different worktrees very annoying to find them and open my editor in the right spot. we'll ship the worktree feature out from the flag next week → tweet link
@levelsio · 2026-05-23T09:24
I don't write code anymore. I haven't written code in I think 6 months? I think everyone is like this no? → tweet link
@levelsio · 2026-05-23T11:32
Every bug fix or new feature on any of my sites I now built live on my VPS, in production, without any staging. Claude Code only failed me 2x in 12 months, it made a small bug and the site was down for 2x 5 seconds. It never lost any data. [...] I have never shipped so fast and smoothly in my life → tweet link
@levelsio · 2026-05-23T11:45
I'm starting to think it has to do with my tech stack why I'm only one doing this. Vanilla PHP + Vanilla JS + SQLite is so simple and basic it's hard for AI to fuck up. The more complex your stack the higher odds AI (or you!) will make a fatal bug. Simple works here. → tweet link
@nummanali · 2026-05-23T16:02
Large corporates are hiring full teams to for AI enablement across functions. The job ad mentions experience with tools such as Claude, Codex and OpenClaw 🙂 Paying up to £200K fyi → tweet link
Agent Harness Debate & Ecosystem Dynamics
@kunchenguid · 2026-05-22T23:13
[Thread] "it's in the model companies' business interest to do so" — i have no doubt about this. models are becoming commodity, they need vertical integration, owning the customer relationship etc. my argument is not about why they would do it, but whether it's a good thing for the world [...] I appreciate how OpenAI has allowed 3p harnesses to integrate with the codex subscription and kept codex cli open source since the beginning — this is a lot more friendly to the developer ecosystem → tweet link
@jezell · 2026-05-22T22:46
Seems like this week's Google I/O was one of the most forgettable in recent history. → tweet link
Hermes Agent & Nous Research
@Teknium · 2026-05-22T21:04
Thanks to Julien Grok Build v0.1 now has its appropriate 256K context length in Hermes - sorry bout that! → tweet link
@Teknium · 2026-05-23T02:51
I said I'd update if they follow through, and now, Anthropic has granted us security privileges with Claude. → tweet link
@Teknium · 2026-05-23T02:59
Hey all security people, am looking for comments and suggestions on this PR, if you have constructive thoughts or comments please check the PR out and let me know as we build towards far improved security for secrets → tweet link
@sudoingX · 2026-05-23T05:17
interesting. my timeline is full of nous research badges suddenly. something is being assembled. whatever it is, a liftoff is coming. → tweet link
@Teknium · 2026-05-22T22:20
Hermes on your watch? nicee → tweet link
Open-Source AI Advocacy
@TheAhmadOsman · 2026-05-23T07:13
Opensource AI is not just about models, it is about agency. How? AI is becoming civilizational infrastructure. If intelligence becomes something people can only rent from a few closed institutions, the public loses not only software freedom, but operational freedom: the ability to study, build, repair, deploy, audit, adapt, teach, and preserve intelligence systems outside capitalistic institutional permission → tweet link
@TheAhmadOsman · 2026-05-23T00:20
HELL NO. If you give OpenAI access to your financial accounts you might as well post your passwords in the reply → tweet link
@TheAhmadOsman · 2026-05-23T00:07
DROP EVERYTHING. The bible for how LLMs work is now available online to read FOR FREE [...] Covers Tokens/Tokenizers, Transformers, Attention, KV Cache, Prefill vs Decode, Decoding Controls, Agents/Tools, Fine-tuning, Multimodality — then local AI: Quantization, VRAM Math, Hardware Tiers, Runtimes, Privacy, Benchmarks → tweet link
@TheAhmadOsman · 2026-05-22T20:13
Good example of Performative Inference / Local AI grifting — MiniMax-M2.7 on 4x M5 Max MacBooks (~$22,000 USD). Longest prompt: 338 (???). Max context: ~3k (???). 45 tok/s (lol). Single prompt, no parallel requests → tweet link
OpenClaw / KiloClaw Ecosystem
@RealGeneKim · 2026-05-22T20:12
OMG, I love KiloClaw! [...] It's a hosted OpenClaw service, with superb documentation and walkthroughs. [...] If you've been OpenClaw-curious like me (I've been lurking since December!), but afraid to try, give KiloClaw a try!!! → tweet link
Robotics & Hardware Hacking
@badlogicgames · 2026-05-23T17:26
new electronics project. don't have the time to design my own little AI robot. still waiting for my HF robot. so i'm going to frankenstein this €9 one and give it a brain. wish me reversing luck! → tweet link
@sudoingX · 2026-05-23T05:30
one more dgx spark would cure me. → tweet link
AI Image Generation
@LinusEkenstam · 2026-05-23T16:36
steal this prompt: "Surreal mid-century editorial illustration of an eccentric version of the person in the photo and a Zebra [any animal] inside a stylish bar [any location]..." — works with ChatGPT Image 2, Nanobanana Pro → tweet link
Industry Trends & Big Tech Restructuring
@levelsio · 2026-05-23T11:52
Also I'm starting to see big tech companies start to operate like me now. They're firing HR, middle management, merging roles. Everyone becomes a builder with AI. I think because with AI people like me have access to teams of 10,000 employees but in the form of AI. So the difference between a solo builder with AI and a big tech company with 10,000 employees is smaller than ever! → tweet link
@levelsio · 2026-05-23T14:42
You now have non-tech normal people outship tech people in terms of reaching revenue fast [...] an Indonesian girl, who's tapped into TikTok culture, knows what to ship, can't even code but ships it fast thanks to AI and gets to $800 MRR in the first month. [...] There is little to any benefit being in tech now over normal people → tweet link
PufferLib & ML Research
@jsuarez · 2026-05-22T23:06
Core optimization improvements to PufferLib today: MinGRU h x 3h projection layer -> orthogonalize the 3 slices separately in Muon / Replace NS with Polar Express / mup scaling makes it easier on our sweeps to tune learning rate jointly with model size / Aurora update on rectangular matrices → tweet link
Project Releases & Dev Updates
@badlogicgames · 2026-05-23T10:09
People of [project]. Minimal read release! Big shoutout to Armin Ronacher, aka @mitsuhiko, creator of Pi, without whom none of this would be possible. ❤️ → tweet link
@badlogicgames · 2026-05-23T09:08
People of [project]. I want to make this the new default. No setting. Not much value in read showing the first X lines, as long as we still show offset/limit if given by the model. Mo minimal, mo better. → tweet link
@steipete · 2026-05-22T22:06
We used bots so far to enforce a 10PR per "person" limit. Great to see GitHub shipping that natively! → tweet link