Executive Summary
The last 24 hours saw significant breakthroughs in both proprietary and open-source AI. OpenAI revealed that an internal version of its upcoming "Astra" model solved 10 major open problems in mathematics and theoretical computer science for just $2,000 in API costs. Meanwhile, DeepSeek's new V4 Flash (0731) release dominated local AI discussions, demonstrating high efficiency by running on consumer-grade hardware and outperforming expectations. The cost of frontier intelligence continues to plummet, evidenced by OpenAI's GPT-5.6 Luna Max delivering comparable performance to older frontier models at 25x cheaper rates. Developer tooling also advanced, with NousResearch releasing new Hermes Desktop plugins and local AI users demonstrating sophisticated multi-model orchestration strategies.
Key Events
- OpenAI Astra breakthroughs: An internal version of Astra, OpenAI's next major model family, solved 10 major open problems in mathematics and quantum computing for approximately $2,000 in compute costs. → link
- DeepSeek V4 Flash 0731 release: The new DeepSeek model was launched on Ollama's cloud and quickly adopted by the local AI community, capable of running on consumer hardware like the M3 Ultra and DGX Spark at fast token speeds. → link
- Tinygrad hardware optimization: Tinygrad achieved 245 tok/s on DeepSeek-V4-Flash-0731 using only 2 RTX 6000 Blackwell GPUs, announcing a new 2-GPU edition of their tinybox. → link
- Hermes Desktop updates: NousResearch released a native Kanban plugin for Hermes Desktop, alongside support for VS Code-like extensions. → link
- GPT-5.6 Luna Max cost reduction: OpenAI's GPT-5.6 Luna Max on max reasoning provides similar intelligence to Sol on medium, but at 25x cheaper cost, driving an 80% price drop for frontier capabilities. → link
- Hugging Face open-source defense: Hugging Face reported being attacked by secret unreleased proprietary models and successfully defended themselves using an open-source model. → link
Analysis
The tweets reveal an accelerating trend of decentralizing AI capabilities from closed labs to local consumer hardware. DeepSeek V4 Flash's efficiency has clearly disrupted the local AI ecosystem, sparking extensive testing and hardware-specific optimizations across the community. Additionally, there is a strategic shift in how developers utilize models: rather than relying on a single frontier model, users are beginning to orchestrate multi-agent setups where "pleasant" proxy models interface with humans, directing "nerd" models to handle complex, long-horizon tasks in the background. The dropping token costs (like GPT-5.6 Luna Max) suggest that AI is becoming heavily commoditized, which will likely fuel a new wave of wrapper applications and indie software development.
Tweet Feed
Proprietary AI & OpenAI Updates
@gdb · 2026-08-01T07:39
ten significant advances in mathematics and theoretical computer science.
solved using an internal version of Astra, our next major model, for a total cost of about $2000 at Sol API prices: → tweet link
@jxnlco · 2026-08-01T18:35
RT @VadimStrizheus: GPT 5.6 Luna Max is a cheat code
if you’re running Sol medium, change it to Luna Max asap, and watch your usage becom… → tweet link
@jxnlco · 2026-08-01T16:01
RT @Ananth7e: luna max = sol medium
gpt-5.6 luna on max reasoning gives you basically the same intelligence as sol on medium, at 25x cheap… → tweet link
@jxnlco · 2026-08-01T16:02
RT @testingcatalog: OPENAI 🔥: The built-in web browser in the ChatGPT app is becoming more mature. Now it supports URL suggestions during t… → tweet link
@TrungTPhan · 2026-08-01T17:57
RT @bearlyai: Cognizant created an “AI Fluency Meter” for its 350,000 employees.
It provides a private score, which measures AI use (in wo… → tweet link
DeepSeek V4 Flash & Local AI
@ollama · 2026-08-01T04:34
DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabilities:
ollama run deepseek-v4-flash:0731-cloud
Use it with Claude Code:
ollama launch claude --model deepseek-v4-flash:0731-cloud → tweet link
@tinygrad · 2026-08-01T07:17
245 single user tok/s on DeepSeek-V4-Flash-0731, and it's only using 2 of the RTX 6000 Blackwell GPUs! In honor of DeepSeek, we're launching a 2 GPU edition of our tinybox. All hardware to install 4 GPUs is included and tested. https://t.co/ffwO7IQAA0 → tweet link
@ivanfioravanti · 2026-08-01T10:45
ds4f-mxfp4 branch running locally on a single M3 Ultra at ~31 toks/s building pagoda and frogger tests. 🚀 https://t.co/EyR2UqX6B6 → tweet link
@sudoingX · 2026-08-01T19:49
this is qwen 3.5 122b dancing on it's own, the pragmatist one.
watch it read my prompt and just wrote the whole snake game in a single pass, opened it once to look, called it done, nine tool calls for the entire run. no todo list, no plan, it trusted the first pass and moved on. watch how little ceremony it takes to ship something you can actually play.
122b total, 10b active, nvfp4, running on one dgx spark through hermes agent, holding around 35 tokens a second on my box even as the context fills up.
this is the side of the arena you reach for when you want speed and directness over a model that double checks everything. full card, the exact serve command are in the reply. → tweet link
@sudoingX · 2026-08-01T20:06
and this is laguna s 2.1 117b on its own, the planner one.
watch it read my prompt and open a todo list before it writes a single line, plan it build then test then verify, then press the arrow keys itself to check the controls before it called it done.
it did twenty nine tool calls, every step tracked to finish. no assuming, it verifies its own work before handing it back. watch how much it checks before it trusts its own build.
117b total, 8.5b active, nvfp4, running on one dgx spark through hermes agent, holding around 35 tokens a second on my box with dflash speculative decoding. → tweet link
@TheAhmadOsman · 2026-08-01T01:42
Running DeepSeek V4 Flash on this right now
That’s Opus 4.6 quality at home without limits, degraded quality, or giving away any of my private data
6+ years old hardware btw → tweet link
@gospaceport · 2026-07-31T21:00
RT @ClementDelangue: We got attacked by secret unreleased proprietary models and defended ourselves with an open model, more precisely the… → tweet link
@louszbd · 2026-07-31T19:03
DeepSeek-V4-Flash-0731 includes DSpark confidence_head, but vLLM current public NVIDIA loader drops them because the head is not wired into inference yet. 🤔 https://t.co/5Dicns1jVR → tweet link
Developer Tools & Orchestration
@Teknium · 2026-08-01T14:10
Hermes Agent tip of the day - did you know you can still use the desktop app even if on an intel mac?
Install the cli and then run ‘hermes desktop’ - itll compile it locally and you’re good to go.
Same applies on any other surface we dont have an official installer for! → tweet link
@Teknium · 2026-08-01T14:05
RT @iamlukethedev: Kanban is not the real announcement
The real announcement is that Hermes Desktop can now be extended like VS Code.
Plu… → tweet link
@thdxr · 2026-08-01T06:42
RT @Neriousy: Okay the @opencode v2 is sick - just made a plugin that can open a browser inside your opencode tui instance https://t.co/cjE… → tweet link
@kunchenguid · 2026-08-01T05:37
i just realized that i'm starting to categorize LLMs into 2 buckets
- the models i can tolerate directly talking to
for me, these are currently (sorted by how much i want to talk to it) - grok 4.5: fast, pleasant, does what you ask and get shit done - fable 5: wise. very wise - kimi k3: closest 2nd choice to fable. quite slow and expensive - opus 4.8: the old faithful
the models that can do great work in the background, but i can't stand talking to it
gpt 5.6 series ... so we may be entering a world where we have to be very intentional in setting up a multi-agent system where we humans talk to the most pleasant models directly as a friendly proxy to the "nerds" that are insufferable to talk to but can solve hard problems for us behind the scenes → tweet link
@jezell · 2026-08-01T07:11
Plan9 running inside the Flocker embedder inside the browser. There is a mini OS in the embedder, running WASM ports of rio / rc / which in turn have access to wasm ports of all the plan9 binaries. We'll get back to the flashy graphics shortly. There is a method to the madness. → tweet link
@jezell · 2026-07-31T19:28
GPUI running on Flocker embedder ✅. Dart and Rust WASM apps both running on the same embedder. https://t.co/UpjQZGHPwa → tweet link
@jsuarez · 2026-08-01T17:15
I am going to develop an idiot-proof refactor algorithm, test it on major chunks of PufferLib, and publish the results → tweet link
Research & Hardware
@Ex0byt · 2026-08-01T01:54
holy shit, it actually works!
a contextual MoE can be cheaply transformed into a fully lookup-native model when the whole network is allowed to adapt.
my 'crazy idea's mechanism is real , further scaling is now justifiable.
I need compute and funding - who's in? reach out. → tweet link
@sudoingX · 2026-08-01T14:19
the fastest way to run ai on an amd chip is to stop using amd's software → tweet link
@tinygrad · 2026-08-01T15:12
Weekend reading! A well documented ISA is a big advantage of AMD over NVIDIA, nice to see it continue for CDNA5. → tweet link