Executive Summary
The last 24 hours saw major momentum in the open-source and local AI space, highlighted by the imminent release of Qwen 3.8 Max and a local 27B model. MiniMax H3 was publicly released, bringing high-quality text-to-video generation to local hardware setups like DGX Spark and Apple Silicon. Meanwhile, DeepSeek V4 Flash 0731 is proving to be a massive leap for local inference, running effectively on 190GB VRAM systems. Developer tooling also saw significant upgrades, most notably a highly optimized Hermes Agent update and continued exploration of long-running autonomous agents like Grok 4.5 and OpenAI's Codex.
Key Events
- Qwen 3.8 Max and 27B Open Weights Announced: Alibaba's Qwen team confirmed that Qwen 3.8 Max (a 2.4T parameter frontier model) and a new local 27B model will drop next week, significantly boosting the capabilities of consumer-grade hardware. → link
- MiniMax H3 Goes Public: The new MiniMax H3 video generation model is now available, with developers successfully running it locally on DGX Sparks and Mac hardware for high-quality video output. → link
- DeepSeek V4 Flash 0731 Sets Local SOTA: DeepSeek's latest flash model is being hailed as the current state-of-the-art for 190GB VRAM/unified memory systems, running effectively on local hardware like a single DGX Spark. → link
- Hermes Agent Major Update: Nous Research's Hermes Agent received a massive efficiency update, reducing token bloat and wasted turns by tracing 250k conversations, making it dramatically better for smaller and local models. → link
- Three.js Ported to Flocker: Three.js was successfully ported to
three_flocker, enabling buttery smooth cross-platform 3D rendering via WebGPU in a WASM app using a single binary. → link
Analysis
A clear pattern over the last 24 hours is the aggressive push of frontier-level capabilities down to local and edge hardware. Labs like Qwen and DeepSeek are releasing smaller, highly optimized models (e.g., the 27B open weights) that punch far above their size, extending the usable lifespan of consumer GPUs. On the software side, agent harnesses like Hermes Agent are focusing heavily on token efficiency, realizing that context waste is the primary bottleneck for smaller local models. Furthermore, a growing discourse around LLM training highlights a tension between RLHF (human-likability) and RLVR (machine-verifiable rewards), explaining why newer coding and reasoning models feel more robotic but perform better on objective tasks. Watch for the Qwen 3.8 drop next week to trigger a massive wave of local benchmarking and hardware utilization tests.
Tweet Feed
AI Model Releases & Updates
@Teknium · 2026-08-03T04:53
Qwen 3.8 Max and a new local 27B Qwen 3.8 is coming!
Thanks @Alibaba_Qwen for featuring Hermes Agent in the release video! → tweet link
@TheAhmadOsman · 2026-08-03T04:59
Looks like we will be getting the weights of Qwen3.8-Max and Qwen3.8-27B next week
We are eating GOOOOD boys → tweet link
@Ex0byt · 2026-08-03T14:17
incoming Qwen-3.8-Max and Qwen-3.8-27B open weights will be epic. → tweet link
@ivanfioravanti · 2026-08-03T07:05
RT @MiniMax_AI: MiniMax-H3 Is Now Publicly Available → tweet link
@ivanfioravanti · 2026-08-03T06:58
MiniMax H3 download in progress on DGX Spark! Let's try!!!! → tweet link
@TheAhmadOsman · 2026-08-02T19:25
RT @MikeBradleyAI: TLDR on @deepseek_ai 0731 V4 Flash. It is comfortably the current SOTA for 190GB VRAM or unified memory based systems. → tweet link
@TheAhmadOsman · 2026-08-03T16:16
Yet another thing where Kimi K3 is SoTA and beating the frontier → tweet link
Local AI & Hardware
@sudoingX · 2026-08-03T07:38
one desk box. 128gb. 3 minds living inside it, and none of them think alike.
- deepseek v4 flash is the autistic philosopher in the corner room...
- laguna s 2.1 is adhd but medicated...
- qwen 3.5 122b is the sprinter that never fumbles... → tweet link
@sudoingX · 2026-08-03T06:54
qwen just posted frontier numbers with a 2.4t model, then in the same breath said the 27b is going open next week. they are adding years to your hardware. my 3090 class card ran qwen 3.6 27b as the king, now 3.8 shows up at the same size and the same card gets smarter for free. → tweet link
@sudoingX · 2026-08-03T00:05
for everyone who wanted to watch deepseek v4 flash 0731 actually dance on a single dgx spark, here it is. i asked it to add auto aim and twin fire boosters to a space shooter it had already built... → tweet link
@ivanfioravanti · 2026-08-03T16:52
MiniMax H3 960 × 544 - 10 seconds - generated in 16:57 mins on DGX Spark. Time to test on Apple Silicon. Love this Nolan/Inception style videos! → tweet link
@ivanfioravanti · 2026-08-03T09:58
And here it is as promised! A clip created locally on DGX Spark with ComfyUI Image to Video MiniMax H3 template! 36 mins for resolution 896x1184 8 secs length → tweet link
@jezell · 2026-08-03T15:13
RT @ivanburazin: For years, every GPU request I got was for H100s by default. Since last week, every single message has been asking for B20… → tweet link
Developer Tools & Agents
@Teknium · 2026-08-02T23:56
Hermes Agent is now dramatically more efficient, especially for smaller/weaker/local models! With the help of @nvidia's Nemo Relay and several other strategies Hermes was able to identify a ton of optimizations... → tweet link
@Teknium · 2026-08-03T18:25
So much in the Herald Release of Hermes Agent - Voice Activated Chats (Replace your google home and alexas!) - Plugins and Kanban for the Desktop GUI App - Agent2Agent Protocol - Outbound Webhooks → tweet link
@sudoingX · 2026-08-03T07:17
grok 4.5 has been grinding on a single task for over 3 hours straight and it will not quit. i told it "keep pushing until you make it" and walked away. still going. and it's genuinely problem solving, it tried to load an 83gb image model in bf16, hit a memory thrash, and got killed by the oom reaper on signal 9. grok found the runaway process itself, killed it, confirmed the memory freed, then pivoted to int8... → tweet link
@MengTo · 2026-08-03T17:27
It’s incredible that you can let Opus 5 run for two hours in Claude Code and get a cinematic Three.js movie inside one HTML file. → tweet link
@swyx · 2026-08-02T19:08
one way i'm developing Forge is to use it to host all my projects going forward. so I often find myself having to bounce back and forth between platform and product. sharing neat trick - in @openai codex you can @ a thread + queue up the @... → tweet link
@gdb · 2026-08-03T00:29
codex for customer feedback -> roadmap → tweet link
@jack · 2026-08-03T03:39
RT @tlongwell_bzz: Buzz ships with the 3.8MB Buzz Agent. It works with any OpenAI-compatible provider. Tiny footprint, highly effective. → tweet link
@louszbd · 2026-08-03T10:51
Thanks for all the feedback! We’ve fixed many of the issues reported and added some new features:) You can now: - Keep memory scoped to each project. - Create automations in form mode by default, and manage them in scheduled and idle views. → tweet link
AI Research & Training
@kunchenguid · 2026-08-02T19:52
if you've been using latest frontier LLMs, it's almost certain that you would have noticed by now the newer models have become worse to talk to... how did that happen? ...RLHF = training the model to be likable by humans... RLVR = training the model to be accepted by machines... RLVR is more scalable... now you see why the newer models are becoming less and less likable? → tweet link
@sudoingX · 2026-08-03T00:50
your engineers aren't lazy. benchmarking your own workload across model sizes is genuinely brutal... that's why almost nobody does it, and it's the exact gap that quietly sinks companies. the fix is an internal benchmark that knows every corner of your workload... → tweet link
@thdxr · 2026-08-03T15:58
every company is publishing random ai benchmarks because they're desperately looking for some way to market their product it's useless noise. no one working on the things they're benchmarking pay any attention to them → tweet link
@levelsio · 2026-08-03T11:55
Yongfook posits that whenever ChatGPT or Claude switch to a new model release, the model is trained differently and that could affect mentions and then traffic to your business It's a great point and I never thought of this... Interesting time for AI SEO → tweet link
Software & Graphics
@jezell · 2026-08-03T06:07
@threejs ported to flocker as three_flocker. Skia provided glyph outlines extruded and rendered via WebGPU in a WASM app running in the flocker embedder. supported on web, mobile, desktop all with the same WASM binary. Buttery smooth cross platform 3d. → tweet link
@jezell · 2026-08-03T17:12
Skin and bones working in three_flocker. Party time 🎉. WebGPU + WASM never looked so good. → tweet link
@RydMike · 2026-08-03T17:48
Insane Flutter fork cooking 🔥 → tweet link