Executive Summary
Developers are increasingly shifting focus toward AI model efficiency, advanced agentic workflows, and complex coding assistants like Claude Code and OMP. Local AI tooling and video generation models, including MiniMax H3, LTX 2.5, and Qwen 3.8, are driving heavy compute experiments on Apple Silicon and Nvidia hardware. Meanwhile, humanoid robotics achieved new milestones in agility and coordination, and major tech acquisitions like Stripe's potential acquisition of OpenRouter are shaping the AI infrastructure landscape.
Key Events
- Rumors emerge that Stripe might have acquired OpenRouter to enhance AI model routing and payment integrations → link
- Nvidia notifies customers about AI-related price hikes of over 15% → link
- Chinese humanoid robots achieve major milestones, including a 100-shot tennis rally by Galbot and a 400m race completed in 38.15 seconds → link
- Open-sourcing of "backpass" by @kunchenguid, a tool to synthesize past agent sessions into optimized AGENTS.md files for LLMs → link
- Ollama collaborates with Poolside AI and Nvidia to launch new open models, including Nvidia's Nemotron → link
- @Teknium announces Hermes Desktop updates, moving toward a version-only update/install system with signed compiled binaries → link
- Developers heavily test local video generation models like MiniMax H3 and LTX 2.5 on Apple Silicon (M3 Ultra, M5 Max) → link
Analysis
There is a clear trend toward sophisticated "agentic" workflows, with developers actively debating the best harnesses (OMP, Hermes, Claude Code) and creating standardized instruction files (AGENTS.md/CLAUDE.md). AI coding models are shifting from basic autocomplete to autonomous PR generation, raising new questions about model efficiency and context window management. Additionally, local AI development on Apple Silicon is maturing rapidly, enabling complex video generation and large language model execution outside of traditional cloud environments. Watch for further consolidation in AI routing/payments and open-source model releases in the 70B-210B parameter range.
Tweet Feed
AI Models & Research
@ivanfioravanti · 2026-08-23T18:19
My main AI driver model currently is GLM 5.3! It can handle everything!
Kimi K3 is great, but limits are low even in the Vivace (top) plane. → tweet link
@kunchenguid · 2026-08-23T17:06
oof.. even within anthropic, opus 5 is only selectively used for specific nerdy optimization problems
why? because it’s unusable otherwise. it doesn’t know how to talk to humans
why? because it’s trained by machines checking only whether it passed tests → tweet link
@ivanfioravanti · 2026-08-23T16:15
You wake up in the morning and... reality has changed!
MiniMax H3, before and after. Prompt: Replace the view outside the window in @ Video1 with a view on black hole gargantua
MiniMax H3 Tutorial on Notion is full of great tips & tricks like this one: https://t.co/lVJA7uYZ7L https://t.co/s1dW11FBGr → tweet link
@ivanfioravanti · 2026-08-23T15:21
RT @stevibe: I tested every Unsloth quant of Qwen3.8 27B on canvas coding, and Q8 was NOT the winner 👀
Setup: 2 prompts: > growing tree an… → tweet link
@ivanfioravanti · 2026-08-23T14:51
I'm testing the preview of Primus by @transformerlab to test an idea for a research I had around Qwen 3.8, I think my assumption was wrong, but it's part of the game. I hope the budget will be enough (50$ remaining) 🤞
One last run to go: "The paper's thesis is deliberately still unwritten, because which story it tells depends on the number that has not arrived. Six paragraphs are marked as waiting on it; everything else is ready to draft the moment it lands."
I was not expecting Qwen 3.8 27B to think SO MUCH in xhigh. If to reach 52 of Intelligence Index we have to spend a fortune in tokens... probably we need to find a different solution.
But let's see at the end of this test. 🤞 → tweet link
@kunchenguid · 2026-08-22T20:15
time to reveal the results!
most people believe gpt-5.6-sol is closer to the opus tier
and the distribution of votes was extremely interesting in a few ways -
when i first posted the poll, opus was about 70% out all votes - a very clear winner
as the poll went on, fable got more and more votes
how could that happen? well the only explanation is that there’s a bias created by who’s most active on X, and who’s following me, because those would be the quickest to engage on the poll
my followers skew towards power users, so if i am to take an educated guess, it would be that the more you use these AI models, the more you would think sol is closer to opus tier - of course individual opinions may vary, but seems that’s the skew
in the end, i’m surprised how even the distribution is. it suggests that across a wide range of usage patterns, sol is somewhere in the middle - better than opus, but not quite fable like. do you agree? → tweet link
@TheAhmadOsman · 2026-08-22T23:31
~70B Dense, ~120B MoE, and ~210B MoE classes will get revived this year and labeled SoTA at home → tweet link
@TheAhmadOsman · 2026-08-22T20:39
The Ultiamte Step-By-Step LLM Engineering Projects Roadmap
- Build a tokenizer
- Learn embeddings
- Implement RoPE / ALiBi
- Hand-wire attention
- Build MHA
- Build a Transformer block
- Train a mini-former
- Compare objectives
- Build sampling
- Speculative decoding
- KV cache
- MQA / GQA / MLA
- Long context
- FlashAttention
- Hardware budgets
- Toy MoE
- Sparse model trade-offs
- State-space / linear attention
- Diffusion language models
- Data pipelines
- Synthetic data
- Scaling laws
- SFT / DPO / RLHF / GRPO
- Quantization
- Serving stacks
- Eval harnesses
- RAG
- Tool use / agents
- Vision-language adapters
- Interpretability
- Red-team suite
- Full capstone model system → tweet link
@thdxr · 2026-08-23T02:17
lot of guesses on what ox alpha is but they are all wrong, kinda disappointed so just going to tell you
ox alpha is a new kind of llm that recursively updates a persistent latent state instead of reasoning entirely through tokens
this lets internal representations converge before anything is actually decoded
those attractors generate shards that encode transformations between latent states rather than the states themselves
at sufficient density these shards compose into metaparameters that dynamically alter the residual geometry of the model without changing its weights.
we built it because there was one thing simply too large to fit inside the context window of any existing model
your mom → tweet link
@ivanfioravanti · 2026-08-22T21:08
These LTX 2.5 22B video by Paolo are beautiful and fast with MLX on M3 Max! I need to test it again on M3 Ultra and M5 Max! → tweet link
Developer Tools & Agents
@sqs · 2026-08-23T18:54
Your orb's portals, now on your own domains. A bit nicer when sharing your agent WIP or using them for long-lived apps.
https://t.co/Z4DoTNKMh8 https://t.co/AhTsjolXr7 → tweet link
@MatejKnopp · 2026-08-23T18:28
Flutter VSync cleanup going strong. Can your slop machine do this?
Just kidding. The slop machine is being very helpful. But in a world where everyone is merging 10k line PRs left and right it's nice to be able to take a step back and ask - do we really need this shit? https://t.co/IuoiAL6bdt → tweet link
@TrungTPhan · 2026-08-23T17:22
wow, the English-to-Claude translator really works https://t.co/Kx0hr8br0F → tweet link
@ivanfioravanti · 2026-08-23T17:09
Quick trick to herdr users: if you want to connect to a remote herdr server, without ssh first, just: herdr --remote servername https://t.co/o72sewqmvM → tweet link
@MengTo · 2026-08-23T17:01
This blew up, so I added 60 more three.js components from experiments I made over the last two weeks.
I'm still amazed by how well Claude Code can create variants from source code, so I went wild with new themes, environments, and interactions.
Some three.js prompts I used: - Environment: mountain, city, bridge, desert, forest, space, stars, earth, moon - Colors: sunrise, blue sky, sunset, nighttime, pastel, vibrant, monotone, or based on the environment - Style: minimal, retro, glass, brutalist - Typography: serif, sans, pixel, tall, wide, letter spacing, min 11 px - Animation: orbit, physics, particles, ripples, braille, wisps, zoom, attract, explode - Interactions: scroll, hover, drag, click, reveal when visible
I now have two Claude Max subs and two Codex Pro subs to keep up. I mainly use Opus for front-end and design, and Codex for everything else, including building this site and repo.
This was my fastest launch ever, and it's already paying back all the time I spent creating these crazy experiments.
To the 30 people who bought Pro or Lifetime, thank you. You're making it possible for me to keep spending so much time on this passion project. → tweet link
@jezell · 2026-08-23T14:51
Calling it now GPUI is the next big thing. Too many people building on it. It's inevitable that it's going to blow up. https://t.co/GOt3P5mKp8 → tweet link
@jezell · 2026-08-23T01:16
RT @jdxcode: usage-rs (my blazing fast rust cli framework) is ready for testing! ~2500x faster than clap!
It’s mostly clap-compatible, but… → tweet link
@kunchenguid · 2026-08-23T04:22
many people asked me how to write CLAUDE.md or AGENTS.md, and i see lots of bad advice flying around
so i took some time to write down a guide in https://t.co/v9rrkWKEFr
tl;dr - handwrite your user level AGENTS.md - for project level ones, you don't write it. you train it like a neural net
i also open sourced my private solution "backpass" at https://t.co/DkM9b4TcZ0 - it samples your past agent sessions for a repo, distill key learnings and losses, synthesize them, and produce a gradient descent step as a proposal that you can review and apply to improve your AGENTS.md and project level skills
easiest way to run it is just "npx -y backpass" in your repo
hope it helps! please share with whoever you think can benefit from it → tweet link
@Teknium · 2026-08-23T02:14
Just FYI we are working on moving towards a version-only based update/install system, compiled binaries, etc.
This'll mean more stable, way faster, signed hermes desktop app installs/updates, less friction, and each update is linked to a patch or full version release.
CLI will still pin to main, and update as usual if you only work through that. → tweet link
@Teknium · 2026-08-23T00:22
RT @HermesWatcher: Okay, This is a big Hermes Desktop update 👀
Hermes can now actually use the Preview pane.
Not just open a page. Not ju… → tweet link
@KingBootoshi · 2026-08-22T23:29
OK i'm trying OMP.
it's hard to try new ai tools/harnesses. i've been using default codex/claude code only. but other harnesses seem better?
agh there's too many things to try.
too many harnesses. too many skills. too many mcps. too many JUST TOO MANY THERE'S TOO MANY THINGS → tweet link
@TheAhmadOsman · 2026-08-22T22:35
I use OMP mainly now, and every time I want to repurpose a harness I just clone vanilla Pi and work with my main agent on repurposing it for the project's goals → tweet link
@ivanfioravanti · 2026-08-23T06:17
After having tested tons of combos I decided that I'll stay on: herdr + omp for now. Let's see if I'll be able to become a super master and increase productivity! → tweet link
@jezell · 2026-08-23T06:16
I've been running 7 parallel goals all day today. Let's see if that reset hits before it drains this sub. https://t.co/kgq2goI9Tu → tweet link
@badlogicgames · 2026-08-23T08:51
recommended reading. i did quite a bit of "hard" stuff in my programming life. but i had to pick my battles, because even tho i knew how to do things, they still took a sometimes prohibitively long time to do.
with agents, i can do a lot more hard things, and merely steer based on my knowledge and experience. → tweet link
@badlogicgames · 2026-08-23T08:49
RT @mitsuhiko: Some weekend thoughts on how LLMs change the way we start new projects. https://t.co/D94vVsVd2t → tweet link
@RayFernando1337 · 2026-08-22T19:11
RT @Rasmic: Meet Bezalel
Bezalel is a capability plane for your agents (claude, codex, OC, hermes, etc)
one MCP gives your agent: 💻 comp… → tweet link
@TrungTPhan · 2026-08-22T19:52
RT @bearlyai: [NEW] The Bearly AI app now has a fully-feauture web browser.
Users can conduct research with multiple browser tabs using th… → tweet link
Hardware & Local Compute
@ivanfioravanti · 2026-08-23T13:47
I really hope @Apple will allow a way to manually set CPU/GPU frequency in their Apple Silicon chips. High Power is too much, Low Power is too little.
We'd like to be able to adjust this based on our needs. → tweet link
@ivanfioravanti · 2026-08-23T13:32
Here is my 1st Experiment with h3.c on M3 Ultra. You can find the full command + prompt 👇
Total time: 2:42 hours 👀 for: - 1344x768 - 10 secs - 20 steps - reuse 2: let the model do every other denoising step, and extrapolate the rest from the last two predictions. Half the compute. - model: https://t.co/AB8X3Yi4pV - prompt heavily inspired by @lepadphone that creates incredible videos with MiniMax H3! → tweet link
@ivanfioravanti · 2026-08-23T06:47
Today is a busy day, the plan is: - test h3.c by @antirez
- test vllm.cpp by @mudler_it 🚀 - more experiments on https://t.co/60VTRDabpM by @eigenlabs - testing new models on my repo that turns the Hacker News front page into local AI art (love e-ink!) - complete a research running on Primus by @transformerlab I hope my assumptions on Qwen 3.8 27B were right 🤞🏻, research is not cheap, I've already spent quite a lot of money and Qwen 3.8 thinks way too much! - keep testing Lux in Tenebris generation with Hermes + Local AI for everything, before @NTTLuke share it in the open!Let's see if I can share some interesting results.
Links: - h3.c: https://t.co/IY6GiEdQZx - vllm.cpp: https://t.co/eXHLkovs40 - hn_local_image: https://t.co/gPQZwoK7wo - mlx fast: https://t.co/bo3RJEr6hF - Primus: https://t.co/8gCHAWUo5E - Lux In Tenebris: https://t.co/DJOP7SIV6Z → tweet link
@ollama · 2026-08-23T01:59
.@poolsideai engineers were awesome to collaborate with to launch open models.
Excited for what’s to come to open models, and the @nvidia team working on Nemotron. → tweet link
@ivanfioravanti · 2026-08-22T19:47
Here is my first experiment with Phosphene 4.6 and MiniMax H3 on M5 Max: 15 secs 1344x768 video rendered in 31 mins only! 9 steps here, now I'll try to push quality up! https://t.co/OfLObqMuGE → tweet link
Robotics
@TrungTPhan · 2026-08-23T15:20
Thought the humanoid robots were smashing into safety cushions because they couldn’t turn corners while sprinting.
Well, one of them just finished the 400m race in 38.15 secs (topping Wayde van Niekerk’s record of 43.03 secs). https://t.co/IVQyYwtllY → tweet link
@TrungTPhan · 2026-08-23T17:27
RT @bearlyai: In terms of coordination, the Galbot Chinese humanoid robot completing a tennis rally of over 100 shots without a mistake is… → tweet link
Industry News & General IT
@jezell · 2026-08-23T17:50
RT @thsottiaux: 2026 is the year companies start seriously caring about model efficiency and reliability as it becomes critical infrastruct… → tweet link
@jezell · 2026-08-23T17:27
@stevendcoffey seems like Responses API doesn't have same protocol auth option that the Realtime API has? Shouldn't you be able to connect directly to Responses API websockets in a browser? → tweet link
@TheAhmadOsman · 2026-08-23T04:08
My tokens usage for the past month or so
Little surprised, it is more than I expected for sure https://t.co/za3ewDIjBm → tweet link
@tinygrad · 2026-08-22T20:35
Agree that AI can make people faster at things they could already do. If you are going to work on a tinygrad bounty, you can use AI to solve it faster, but if you don't think you could solve that bounty yourself, you are wasting everyone's time trying to solve it with AI. https://t.co/589sSDfemw → tweet link
@jezell · 2026-08-22T20:03
RT @filpizlo: Why am I building SaRCAsm (safe runtime capability enforcing assembler)?
It’s a way to dramatically increase the guarantees… → tweet link