Executive Summary
The tech and AI community focused heavily on agentic coding tools, local inference setups, and fast model performance. New AI models and features like DeepSeek-V4.1-Flash, Qwen 3.8 Flash Next, and OpenAI's Codex updates dominated discussions, with developers running impressive benchmarks across Cerebras, DGX Spark, and Apple Silicon hardware. Developer tooling saw major releases, including Serverpod 4 and Transformers.js v4.3.0, while the broader discourse touched on the economics of AI inference and the shifting market dynamics between frontier labs.
Key Events
- Teknium details how Hermes Agent powered almost 1,400 subagents over 19 hours to refactor 400,000 lines of code, launching a new plugins catalog. → link
- sudoingX shares a comprehensive 8-point guide for setting up robust local agentic systems using Linux, Tailscale, llama.cpp, and tmux. → link
- victormustar announces DeepSeek-V4.1-Flash as the new default model for HuggingChat, integrated with Exa web search. → link
- Ivan Fioravanti breaks the 30 tokens-per-second barrier for DeepSeek V4.1 Flash running locally on a single M3 Ultra Mac. → link
- OpenRouter data reveals that users spent more on OpenAI models than Anthropic models last week for the first time in over 2.5 years. → link
- MengTo open-sources a scrolling virtual-tour landing page along with 8 new three.js skills. → link
- Serverpod 4 is released, described as built from the ground up for agentic coding with full-stack hot reloading. → link
Analysis
There is a clear trend toward empowering developers with highly efficient local AI infrastructure, as showcased by robust setups using DGX Sparks and Apple Silicon. Agentic coding workflows are becoming deeply integrated into standard development pipelines, with tools like Hermes Agent and OpenCode demonstrating powerful large-scale refactoring and multi-agent task delegation. Meanwhile, the shift in OpenRouter spending highlights a cyclical competition between frontier labs, while the developer community continues to demand practical speed, utility, and outcome-based economics over raw benchmark scores.
Tweet Feed
AI Models & Performance
@victormustar · 2026-09-16T11:14
HuggingChat goes whale 🐋 New default model: DeepSeek-V4.1-Flash. Go try it to get how good open AI is now. All you need is a Hugging Face (free) account. Bonus: it searches and crawls the web for you using Exa https://t.co/LfSWSE530c → tweet link
@ivanfioravanti · 2026-09-15T23:39
DeepSeek V4.1 Flash on DwarfStar: the 30 tps barrier has been broken! And now the single M3 Ultra is as fast as two in decode and faster in prefill! Branch is here, give it a try: https://t.co/vzfR7hXSpV → tweet link
@MilksandMatcha · 2026-09-16T16:32
I have tracked my macros (mostly protein) for years now. It is time consuming and extremely annoying. All the apps suck and it take forever.
So I gave 3 models the same food photo and asked for the macros:
Qwen 3.8 27B on @cerebras // 2.46s GPT Astra by @OpenAI // 19.55s Fable 5.1 by @AnthropicAI // 14.74s
All three landed around 1,200–1,300 calories. Qwen finished nearly 8x faster than Astra in this run.
I just want to log my lunch and eat it :) → tweet link
@sudoingX · 2026-09-16T01:17
qwen 3.8 flash next built this landing page in 29 minutes and served it on my tailnet on its own. i am running official fp8 on 2x dgx spark, full 256k context loaded, 45 tok/s with mtp on. i have run deepseek and glm on these boxes and it was good, this one just feels right. qwen 3.8 flash next is multimodal native, a vision encoder in the same weights, so the serve that wrote this page can read the screenshot it took of it. and it stays sharp at the depth where the others start drifting. qwen 3.8 flash next it is. → tweet link
@RayFernando1337 · 2026-09-16T14:31
I can’t gatekeep this anymore. 2 Sparks will change your life. GLM-5.3 Flash EXL3 from Mia is the best out there right now. → tweet link
@RayFernando1337 · 2026-09-16T13:48
RT @MiaAI_lab: Grok 4.6 on High is better than Opus 5 on Max This isn't a guess, it's a simple fact. After 48h of using Grok on the exact… → tweet link
@nummanali · 2026-09-15T19:26
Literally the best paper name Never Give Up This is basically /goal mode for post training → tweet link
@alexinexxx · 2026-09-16T00:48
we should stop publishing benchmark scores and just make model authors sit in a room while people try their model for 10 minutes → tweet link
Agentic Coding & Developer Tools
@Teknium · 2026-09-15T22:59
My first blog - Hermes powered almost 1400 subagents over 19 hours to refactor 400,000 LOC out of Hermes Agent's repo. Read the full story below! → tweet link
@Teknium · 2026-09-16T16:56
Introducing the Hermes Agent plugins catalog! You can access it in your Hermes Desktop app's capabilities section to easily search, browse, and explore community and official plugins! Submit your own as well! → tweet link
@thdxr · 2026-09-16T18:59
oh cool i just remembered we had subagent continuation so i got the review and it had a good idea so i was able to tell it to continue with the implementation in that subagent https://t.co/8N8IF5JGEb → tweet link
@thdxr · 2026-09-16T15:53
in the latest opencode2 you can optionally mention a different model to use for subagents very useful, i use it to get reviews from smarter models https://t.co/iMQMhtd312 → tweet link
@thdxr · 2026-09-16T14:15
OpenCode found potential issues in Search for Extraterrestrial Intelligence software BLISS scans radio telescope data to look for extraterrestrial signals a simple bug could cause them to be missed PR below → tweet link
@jezell · 2026-09-16T16:09
RT @testingcatalog: OpenAI is preparing a Codex Replay feature that lets users test task execution from any imported conversation thread.… → tweet link
@jxnlco · 2026-09-16T16:30
RT @unitygames: We heard you loud and clear 🔊 Introducing the official Unity plugin for @OpenAI’s Codex - a first-party integration availa… → tweet link
@RydMike · 2026-09-16T18:09
RT @ServerpodDev: Serverpod 4 is out. 🚀 Our most ambitious release to date. Built from the ground up for agentic coding. Full-stack hot rel… → tweet link
@sqs · 2026-09-16T18:38
Outcome-based pricing for code → tweet link
@sqs · 2026-09-16T18:08
⌘P (or ^P on Linux/Windows) is now switch-thread for everyone Join me in thanking @homborg for the good idea https://t.co/vhGS6TDQcC → tweet link
@KingBootoshi · 2026-09-16T02:06
experimenting with a dynamic mac os app that pops up whenever my agents have design questions and lets me pick on the spot basically my DX hyperfixation with AI right now is: "how can I make traversing information faster for myself and my agents" i am the bottle neck. my ability to inject my creativity into agents are my bottle neck. the faster i can make this the faster & higher quality my agent builds are → tweet link
@KingBootoshi · 2026-09-16T01:18
i've started naming chats and associating 1 chat per project i shit you not i genuinely just have 1 mega thread per project and i've never been so efficient before i am PRO one mega thread continuously compacting (especially with any harness that is NOT claude code) → tweet link
Local AI & Infrastructure
@sudoingX · 2026-09-16T00:07
if you are starting with agentic systems this month, cloud or local, read this before you pick a model. the setup decides whether you keep your flow, the model only decides how smart the flow is, and almost everyone gets the order backwards. the foundation, running agents across a laptop, multiple desk nodes and a phone: [details Linux, Tailscale, llama.cpp, tmux, and git] → tweet link
@sudoingX · 2026-09-16T03:22
2x dgx sparks on a desk is the endgame and i will die on this hill. i had the frontier in my room this week, with full 256k context loaded, and 37 watts on the gpu readout while an agent clank. under $10k for the pair and the cable. that is a used car, and the used car does not think. if you have been waiting for the local ai moment to be real, this is it, and it is already sitting on my desk. → tweet link
@sudoingX · 2026-09-16T15:54
i got annoyed enough at nvtop showing N/A for memory on both my dgx sparks that i went and fixed it. the fix sums the per-process nvml allocations, which do work on GB10, and reports total as that plus MemAvailable, so it is the memory the gpu can actually reach. left pane patched, right pane stock, 98.690 GiB against an empty bracket. PR open upstream: https://t.co/2ono5ZqO69 → tweet link
@MilksandMatcha · 2026-09-15T20:14
During an outage, every second matters. That is one of the reasons OpenAI is using @cerebras internally. @seanlie describes how fast inference is being used for incident response and critical research, where more reasoning within the available time can improve the response. → tweet link
Open Source & Web Development
@victormustar · 2026-09-16T08:18
RT @nicodotdev: Transformers.js v4.3.0 is out 🎉 And with it our first sub-package: huggingface/transformers-structured-output Force any m… → tweet link
@victormustar · 2026-09-16T13:30
RT @harshagundal: On HF spaces now, up to 70x speed ups on GPUs… enjoy :) https://t.co/lxBNUUdMTL → tweet link
@MengTo · 2026-09-16T16:06
I open-sourced this scrolling virtual-tour landing page, along with 8 new three.js skills. They cover room-to-room tours controlled by scrolling, sky rays, sky backgrounds, falling leaves, four seasons, high-resolution textures, high-poly models, and retina rendering. Live site: https://t.co/YN0dZ8DEVH Source: https://t.co/bHUdat642y Skills: https://t.co/gMuw2gt9FJ → tweet link
@jezell · 2026-09-15T22:02
RT @rauchg: Safari 27 supports JSPI, a magic WebAssembly capability for sync native code to suspend on an async 𝙿𝚛𝚘𝚖𝚒𝚜𝚎. On https://t.co/3… → tweet link
@gospaceport · 2026-09-16T18:30
RT @Blackwellboy: Built Hugging Face Vanished as a just in case monitor. If models ever start getting gated, removed, censored or changed,… → tweet link
@jezell · 2026-09-16T01:40
RocksDB? Bruh. Someone tell those agents about SlateDB. → tweet link
Industry Commentary & Economics
@jezell · 2026-09-15T19:11
RT @OpenRouter: OpenRouter users spent more on OpenAI models than on Anthropic models last week. This hasn't happened for more than 2.5 ye… → tweet link
@MilksandMatcha · 2026-09-16T02:33
i had such a great starbucls session with @kimmonismus today about how philosophy and existentialism can help us think about AI. For a growing set of tasks, having a human in the loop is becoming a choice. Models can do more of the work on their own, and sometimes our intervention just slows them down. → tweet link
@MilksandMatcha · 2026-09-15T23:26
An excerpt from @SemiAnalysis_ on the market for fast inference. 'the world has revealed their preference for fast tokens with their wallets' → tweet link
@swyx · 2026-09-15T19:04
RT @latentspacepod: 🆕 Humanity’s Last Invention — with @RichardSocher! https://t.co/qOKd1AWBLp @recursive_si is the latest neolab to burs… → tweet link
@hnasr · 2026-09-16T14:04
Caching is a cop-out Cache stores such as Memcached and Redis are often used to speed up slow database queries or avoid recomputing expensive HTTP responses. Caching is critical for scalability, but it should not become a crutch... → tweet link