Executive Summary
The past 24 hours saw significant AI market disruption as OpenAI slashed API prices, notably cutting GPT-5.6 Luna by 80% and launching a Cerebras-powered Fast mode for GPT-5.6 Sol to achieve 750 tokens/sec. Concurrently, local AI and open-source hardware optimization are accelerating rapidly, with developers successfully running 122B parameter models on prosumer hardware like the NVIDIA DGX Spark and tinygrad's tinybox using NVFP4 quantization. Developer tools are shifting heavily toward AI agents and code generation platforms, though concerns are rising about AI-generated code eroding long-term code quality. Finally, hardware supply chain dynamics are shifting, with TSMC packaging constraints creating a rare opening for Intel in the AI chip foundry space.
Key Events
- OpenAI announces massive API price cuts (80% for Luna, 20% for Terra) and a new Cerebras-backed Fast mode for GPT-5.6 Sol. → link
- Thinking Machines releases Inkling Small (12B active / 276B total params) with open weights and NVFP4 support. → link
- tinygrad benchmarks GLM-5.2 (120 tok/s) and Kimi K3 (42 tok/s) running locally on AMD tinybox hardware. → link
- TSMC's packaging constraints are limiting AI chip growth, creating a rare opportunity for Intel to rebuild its foundry business. → link
- Ollama announces a partnership with Intel to enable open models on Intel Core Ultra Series 3 processors. → link
- Amp announces its first dedicated "Amp Labs" in Sydney, partnering with Westpac to push AI software development in banking. → link
Analysis
The primary trend is a two-front war on AI inference costs. On the cloud side, OpenAI is aggressively cutting prices to fend off model routers and alternative providers like Cerebras, pushing towards "intelligence too cheap to meter." On the local side, the community is proving that massive MoE models (like Qwen 3.5 122B and Laguna S 2.1) can run efficiently on single prosumer boxes (DGX Spark) using NVFP4 quantization and speculative decoding.
However, as AI-assisted coding accelerates via tools like Codex, Cursor, and Hermes, experienced developers are flagging a de-escalation in code quality. Subtle technical debt is being introduced by AI models applying superficial fixes that pass automated tests but ignore underlying architectural flaws. Watch next for how Apple Silicon Metal kernels and NVFP4 optimizations mature, and whether enterprise adoption of "vibe coding" leads to systemic software maintenance issues.
Tweet Feed
AI Models & Releases
@victormustar · 2026-07-30T18:55
RT @mervenoyann: Thinking Machines released Inkling Small (🦖) + NVFP4
12B active 276B total params, the model performs better than larger… → tweet link
@victormustar · 2026-07-30T18:46
RT @miramurati: Inkling-Small is comparable to Inkling at a quarter the size. Weights are open, fine-tunable on Tinker today. Look forward… → tweet link
@sama · 2026-07-30T17:27
major price cuts today:
80% drop for GPT-5.6 Luna, now $0.20 per million input tokens and $1.20 per million output 20% drop for GPT-5.6 Terra, to $2/$12 *GPT-5.6 Sol gets Fast mode in the API, up to 2.5x the speed for 2x the price, same intelligence https://t.co/erC6u4VoDR → tweet link
@gdb · 2026-07-30T17:47
we've cut prices on luna by 80%, making it by far the most price-efficient model in its class.
a lot of our research is about how to create incredibly efficient models for any given level of intelligence.
excited to see what you all do with intelligence too cheap to meter! → tweet link
@kunchenguid · 2026-07-30T17:36
gpt-5.6-luna price reduced by 80% (!!!)
it was already cheap - this will make it feel like almost unlimited usage
i showed how i built an entire full stack app with luna alone in my recent video https://t.co/Zc0FOGRwPF, and it only used 4% of my weekly quota - with the 80% cut, it means i can build 100 whole apps every single week
openai is really cooking with efficiency here - massive advantage against competition → tweet link
@jezell · 2026-07-30T15:22
Model routers are trending because enterprises are paying insane rates compared to people running subscriptions. If OpenAI loses to the model routers, they lose control of the race. All they have to do is to stop that is offer the lowest price and the best product. In addition to performance improvements they keep touting, today they'll be rolling out additional capacity vs Cerebras, which should allow them to cut margins if they care to win the race. Let's see if we finally get some decent news for API users today. → tweet link
@jezell · 2026-07-30T14:51
RT @goodworse: GPT-5.6 Sol launches on Cerebras TODAY
70 t/s => 750 t/s
the model will be ~10 times faster
GPT-5.6 can now complete the… → tweet link
@jezell · 2026-07-30T14:42
RT @btibor91: In case you missed the drama around shared Claude conversations appearing in search results, Anthropic has just added "noinde… → tweet link
@ivanfioravanti · 2026-07-30T13:17
RT @jun_song: You are closing eyes on Minimax.
You don’t know what’s coming. → tweet link
@thdxr · 2026-07-30T04:06
kimis license around revenue sharing is totally reasonable
the price fixing part is not - you need to get their approval before lowering prices → tweet link
@thdxr · 2026-07-30T05:23
right now i'm very unsure about this next gen size of model
they're so expensive and slow and it doesn't feel worth it for most things → tweet link
Hardware & Local Inference
@sudoingX · 2026-07-30T18:39
did you know? the dgx spark has 128gb of unified memory. that's more than an h100, and that one number is the whole thing people miss when they write it off as a cute little dev box.
the chart is a 122 billion parameter model sitting whole on this box, 113 gigs used, with still enough memory left to serve fourteen people a full 128k context at the same time. an 80gb h100 can't hold this with any real context to work in. a datacenter card that costs more than a car runs out of memory right where this desk side box is still serving a crowd.
the dgx spark is a memory monster that happens to be quiet. it trades top end speed for the ability to hold the models the expensive cards can't, and holding the model whole is the entire game when you want to own it instead of rent it. → tweet link
@sudoingX · 2026-07-30T17:37
a 122 billion parameter model, reasoning through problems and writing code, on a box sitting on my desk pulling 32 watts on load. that's about two light bulbs.
i loaded qwen 3.5 122b onto a single dgx spark in nvfp4 and measured all of it so you don't have to, baseline and with the mtp drafter switched on. the whole picture's in one frame.
raw decode is 28tok/s. flip on mtp spec-decode and it climbs to 35 on a normal task and up to 54tok/s on predictable code, a little drafter model guessing ahead and getting it right 92 to 96 percent of the time. the weights eat 113 of the box's 121 gigs and there's still enough kv cache left over to serve fourteen people a full 128k context at the same time, off a single box, which is a sentence i keep rereading.
next i'm putting this qwen 3.5 on real autonomous builds through hermes agent and running it straight at laguna s 2.1 to see which one actually ships working code. the fight just moved to the part that matters. → tweet link
@tinygrad · 2026-07-30T16:38
Bringing up our first tinybox pro v2 black! This machine is for sale for $160k in our shop if you feel the FOMO and want one. Shipping in 2-4 weeks. GLM-5.2 benchmarks in this 🧵 https://t.co/d9Zcz81imt → tweet link
@jezell · 2026-07-30T15:39
RT @theinformation: TSMC’s packaging constraints are limiting customers’ AI chip growth, creating a rare opportunity for Intel to rebuild i… → tweet link
@tinygrad · 2026-07-30T15:46
We have a 120 tok/s GLM-5.2 and a 42 tok/s Kimi K3 running locally on AMD boxes and it's legit a frontier setup for ~$600k. Will be interesting to see how that dollar amount varies over time. → tweet link
@ollama · 2026-07-30T05:42
Excited to work with the @IntelBusiness team on enabling open models with Intel Core Ultra Series 3. → tweet link
@ivanfioravanti · 2026-07-30T04:51
I was not expecting this jump in performance in just 2 days of https://t.co/qxnebgWQKw and it's now 83.6%! Apple Silicon Metal Kernels are getting some love! → tweet link
@ivanfioravanti · 2026-07-30T04:36
RT @kernelpool: Benchmarks for the Kimi K3 2bit quant thanks to @ivanfioravanti's llm_context_benchmarks tool. Performance holds up nicely… → tweet link
@ivanfioravanti · 2026-07-30T04:29
Decoding speed of Kimi K3 on 2 x M3 Ultra 512GB is keeping up really well with larger contexts! Great architecture behind the scenes. And amazing job gy @kernelpool 🚀 → tweet link
@sudoingX · 2026-07-30T03:41
a huge chunk of people still argue you need a $2000 gpu to do local ai. meanwhile a used 12gb card is quietly doing it on some builder's desk across town.
i've stopped trying to win that one with words. nobody gets talked into local ai, the first time they load a model onto a card they already own and watch it actually reason. that's the moment the argument just ends.
12 gigs is enough to feel it in 2026. quit debating and go stand on the floor, it's way lower and way better than you've been told. → tweet link
@sudoingX · 2026-07-30T02:08
i posted this a couple weeks ago about needing a second dgx spark, and nvidia just shipped me one. the unit and the connectx cable, on a truck to bangkok within a week of me saying it out loud here on x.
sit with that for a second anon. the biggest company on earth moved that fast for one guy in a room in bangkok. no stanford lab connection behind me, no committee, no comms team drafting my posts, just a builder putting real numbers in the open where anyone can tear them apart. the polished accounts with the pedigrees have been saying "democratize ai" for years, nvidia shipped the box to the guy actually doing it. giants aren't supposed to move like this, and that's exactly why the open side feels different right now.
nvidia is carrying real weight for the local ai side, backing the people who own their compute instead of renting cognition from an api, and doing it quietly, on the ground, for whoever's actually doing the work. i think that's the whole reason this movement has legs man.
here's what happens next. two dgx sparks, one connectx cable, 256 gigs of unified memory, tensor parallelism across both. the models that never fit on one box, glm 5.2, nemotron, deepseek v4 flash at a million tokens, they gonna clank now. i'm going to learn a stupid amount from this setup man, and share all of it, same as always.
biggest shoutout from deepest depth of my heart to @Coolmark482 for making all this happen and for seeing the local ai community as worth betting on. thank you. → tweet link
@ivanfioravanti · 2026-07-29T21:09
RT @MiaAI_lab: Qwen3.6-35b-NVFP4 on your @NVIDIAAI DGX Spark has been upgraded✨
95 tok/s single session 317 tok/s at 8 sessions
256K cont… → tweet link
@sudoingX · 2026-07-29T19:17
this morning i called laguna s 2.1 the king of the single dgx spark, but that win came against a model that couldn't even turn on, so i brought it a real opponent. laguna s 2.1 VS the.. qwen 3.5. 122 billion parameters moe.
a few months old now, nvfp4 on same dgx spark, same vllm, both models running their own spec decoder so nobody got a handicap. harness is hermes agent, both fighters loaded now.
on raw speed with no drafter qwen 3.5 is just faster, 28tok/s to laguna s 2.1 20tok/s. switch the spec decoders on and the code fight becomes a dead heat, both land right around 35tok/s, laguna even edges its best run to 37tok/s. but ask them to write plain english instead of code and it splits wide open, qwen holds a flat 35tok/s while laguna drops to 17tok/s.
and the part that stopped me for a sec is laguna's dflash drafter on freeform is actually slower than laguna running no drafter at all. the draft overhead costs more than it saves the moment the text stops being predictable, a speed feature that turns into a tax.
under load it isn't close either, sixteen streams at once and qwen is pushing 226 tokens a second while laguna flatlines at 108. a drafter is a single agent tool, it stops paying the second the box is serving a crowd.
laguna keeps its corner though, it prefills faster, it's smaller on disk, it holds a full million tokens of context to qwen's 262k, and on pure code with the drafter it's dead even. genuinely great coding model. it's just not the untouchable king the knockout made it look.
a model most people already scrolled past, months old, fits a box on your desk and trades punches with the reigning champ. the hardware didn't change. the field just got deeper than anyone's measuring. → tweet link
Developer Tools & Agents
@swyx · 2026-07-30T17:49
RT @latentspacepod: AI engineers are rediscovering ontologies as a way to keep probabilistic agents inside deterministic boundaries. We del… → tweet link
@swyx · 2026-07-30T17:35
RT @aiDotEngineer: 🆕 First Steps Toward Automated AI Research
https://t.co/04QvCDdUCC
Humanity advances by trying things, finding the sh… → tweet link
@jezell · 2026-07-30T17:32
RT @huacnlee: An example of embed WebView to GPUI window.
With this PR change, we can render GPUI elements top on the WebView.
However, t… → tweet link
@RayFernando1337 · 2026-07-30T17:23
Build iOS apps in the cloud! → tweet link
@MilksandMatcha · 2026-07-30T17:23
AI should reduce workflow friction, not add another chatbot.
@sarahmsachs (Head of AI @ Notion) on moving from isolated assistance to durable agent workflows built around a shared system of record.
From the Token Billionaires Lounge, presented by @cerebras and @aiDotEngineer. https://t.co/v0sbffYt2s → tweet link
@gdb · 2026-07-30T16:21
try asking codex to go beyond what you think possible → tweet link
@jezell · 2026-07-30T16:21
750 tokens/sec is super cool, but most of my time is spent waiting for compiles / verification already. People gonna be familiarizing themselves with Amdahl again? → tweet link
@jezell · 2026-07-30T15:35
RT @huacnlee: With A11y support, GPUI can now be fully automated for dev and UI test by Codex.
In this video, I'm trying to solve render l… → tweet link
@RayFernando1337 · 2026-07-30T15:05
As a person who likes to yap out a large plan for very ambitious tasks, I'm excited to give Wayfinder a look and see if it can help me decompose my plan for agents to tackle the hardest problems. → tweet link
@RayFernando1337 · 2026-07-30T13:58
RT @pbbakkum: To set this up: - Get ChatGPT Desktop with a paid account - Add a remote connection (Settings -> Connections -> Add Device) -… → tweet link
@jezell · 2026-07-30T15:40
RT @letstri: Introducing Druk.
A code editor that can be run in any terminal.
Supports both keyboard and mouse control.
Built on the @op… → tweet link
@victormustar · 2026-07-30T12:59
RT @HuggingApps: ABot World 0.5B is out! we have now world models at home, running in real-time on consumer GPU
feed it an initial image,… → tweet link
@victormustar · 2026-07-30T12:59
RT @HuggingApps: StatePlay is out! a framework for game-world models that generate images + game state 🎮
out with a Street Fighter III mod… → tweet link
@Teknium · 2026-07-30T04:02
RT @Mattmaximo: Hermes just negotiated my Spectrum bill down 58% https://t.co/ZWrfEhR7SD → tweet link
@KingBootoshi · 2026-07-30T03:44
fable but with codex compaction
who’s building this ??? → tweet link
@sqs · 2026-07-30T04:36
Announcing the first Amp Labs, 100% dedicated to @Westpac in Sydney.
Westpac uses Amp heavily, and now we're going deeper with them.
The world-class engineers on the Amp Labs founding team will push the frontier of software dev and AI in the bank alongside Westpac folks on the biggest enterprise/banking challenges.
Welcome Gareth, Andrew, Matty, Ryan, and Chris!
The way we've set up Amp Labs is different from anything else out there. It's why such amazing people want to join and why we get so much more done. More to share soon...
https://t.co/EdMPwNe1zR → tweet link
@Teknium · 2026-07-29T22:24
Hermes has deep integrations with @Jack's Buzz now!
Run
hermes updateto get access early before the next release version :) → tweet link
@Teknium · 2026-07-29T19:52
RT @raviojhax: chat, where can I read a deep comparison of openclaw and hermes agent from a dev pov?
who's used them both in daily life an… → tweet link
@sqs · 2026-07-29T23:36
On @theo's stream now to watch the master at work (now that I am a streamer too, I guess). He is the apex godfather of streaming. His Amp mention was a nice bonus. → tweet link
@jxnlco · 2026-07-29T23:52
RT @codestantine: The Codex Visualize skill is awesome. Had no idea it could generate all these kinds of interactive visuals inline.
You… → tweet link
Software Development & AI Impact
@jezell · 2026-07-30T16:29
RT @simonw: I hope there are QA testing experts out there who are thinking "finally, we don't need software developers any more!", rolling… → tweet link
@thdxr · 2026-07-30T13:14
you now have the ability to - play with every possible solution to a problem - refactor everything when you think of better patterns
so many people complaining about the code the LLMs produce, if you're not producing the best software of your life right now something is wrong → tweet link
@thdxr · 2026-07-30T13:04
now everyone gets to complain about how LLMs don't write good code like they're someone who even knows what that is → tweet link
@levelsio · 2026-07-30T11:49
Today @fireship_dev covered this:
Indie hackers may be the first type of developer to go extinct
The classic indie hacking playbook (last decade): learn to code → find a niche → ship a micro-SaaS → build in public → post MRR screenshots after failures → repeat until it works
Coding used to be hard and expensive, so execution mattered far more than ideas
That equation has flipped: execution now costs ~$20/month, so anyone can build personal software instead of paying for SaaS
Example: Why pay $29/month for a product when Claude Code can recreate a better version in 20 minutes?
Building in public on X makes it easy for others to steal your idea/roadmap
AI raises the ceiling (indie hackers become far more productive) but also lowers the floor (the space is now crowded with normies who can vibe-code their own tools)
This trend will accelerate as AI gets cheaper and smarter
Silver lining: the disruption can "fertilize the soil" for those who master distribution, branding, or proprietary data
https://t.co/vyUwc0bsgu → tweet link
@MatejKnopp · 2026-07-29T21:14
There is no way in hell blind use of AI is not eroding code quality.
flutterdev bug report: 32 bit HMONITOR handle with sign bit set will trigger assertion in debug build - the value gets sign-extended to 64bits and then at some point - when sent to Dart - triggers assertion because it's an uint64_t that exceeds Dart signed 64bit integer range.
AI generated PR: Strips sign extension from DisplayId in the win32 embedder - some bit magic, newly added tests pass, all good.
Cool. Except it is the wrong fix. The public embedder API defines the DisplayId like this:
typedef uint64_t FlutterEngineDisplayId;
There is no indication that win32 embedder shouldn't be able to use all bits in the value. The actual problem is in the engine itself where it should convert the displayId to signed before sending it to Dart and thus avoiding the assertion and preserving all bits.
The original fix would work, the assertion would go away, but the codebase would get just a tiny bit shittier and the underlying issue would still be there.
The scary part is that nothing in the PR screams "this is wrong", and on surface it looks sane enough to easily sneak past a human reviewer. → tweet link
@badlogicgames · 2026-07-29T19:57
RT @GeoffreyHuntley: [1] software factories are super real but we need to be realistic. the factory aspects haven’t been cracked yet, the [… → tweet link
@levelsio · 2026-07-29T23:54
RT @robj3d3: SaaS is dead.
Most subscriptions are just a prompt.
Can I Vibecode It? gives you the prompts.
Copy. Paste. Cancel. https://… → tweet link
@MengTo · 2026-07-30T08:26
Animated backgrounds, text effects, buttons, image galleries, and WebGL components. These three libraries cover a lot of what makes a landing page feel alive:
https://t.co/2shwu9xk3o https://t.co/S8qLfsX2rV https://t.co/tECuh9VUFz
Give the components to AI and customize them instead of rebuilding every interaction from scratch.
Landing pages, apps and now even games are becoming ridiculously fun to build. → tweet link
@MilksandMatcha · 2026-07-30T18:23
Rust keeps its ownership rules even across threads. Move transfers data into closures. MPSC channels enable thread-safe message passing. Arc, Mutex + RwLock cover shared state when needed.
Part 4/4 of the Independent Studies series w/ @shreyas4_ at @KernelLabs_ai 🦀 https://t.co/k2mMnDKvqA → tweet link