Executive Summary
The AI landscape is rapidly evolving with Grok 4.5 emerging as a significant competitor to established models, proving SpaceXAI as a third major player alongside Anthropic and OpenAI. Local AI capabilities are advancing with Unsloth's 1-bit Kimi K3 model enabling frontier-level models to run on consumer hardware, while OpenAI released new transcription models and open-sourced the Codex Security CLI. The industry is seeing increased tension between open and closed AI models, with growing concerns about regulatory capture from frontier labs, and a significant shift in hiring patterns favoring AI-native individual contributors over traditional managers.
Key Events
-
Grok 4.5 emerges as major AI model: User analysis indicates Grok 4.5 is the most pleasant frontier model to work with, proving SpaceXAI as a third real player with unique advantages including real-time data access and proprietary hardware infrastructure. → link
-
Kimi K3 1-bit model enables local AI: Unsloth released a quantized version of Kimi K3, shrinking it from 1.56TB to 594GB while retaining ~78.9% accuracy, allowing frontier models to run locally. → link
-
OpenAI releases new transcription models and security CLI: Two new transcription models (GPT-Live-Transcribe and gpt-transcribe) launched alongside the open-sourced Codex Security CLI. → link → link
-
Google DeepMind dismantles AlphaFold team: FT exclusive reveals the dismantling of the team behind the Nobel Prize-winning AI system for predicting protein structures. → link
-
Mitchell Hashimoto starts new company: The co-founder of HashiCorp announced @superlogical, beginning with a terminal multiplexer with a broader vision. → link
-
Opus 5 receives criticism: Users report that Claude Opus 5 flopped, with concerns about training rewarding machine-verifiable outcomes over human-friendly interactions. → link
-
Local AI Hardware Arena benchmarks: Comprehensive benchmarks comparing RTX PRO 6000, DGX Spark, Strix Halo, M5 MacBook Pro, and Cloud for LLM performance. → link
-
Buzz Desktop releases with broad agent support: Jack's open-source workspace adds support for Hermes Agent, custom harnesses, and multiple AI models including Claude and Grok. → link
Analysis
The AI industry is experiencing several parallel shifts. First, the model landscape is becoming more competitive, with SpaceXAI proving it can build frontier models while having unique advantages in real-time data access. Second, there's a growing divide between open and closed AI approaches, with open models like Kimi K3 demonstrating that frontier-level capabilities can be democratized. Third, the industry is seeing increasing concerns about regulatory capture, with frontier labs advocating for slowdowns that critics see as anti-competitive. Fourth, there's a significant shift in developer workflows, with local AI becoming more viable and agent harnesses becoming key differentiators for model training. The hiring market is bifurcating dramatically, with AI-native individual contributors becoming more valuable than traditional managers. Watch for continued pressure on closed model valuations as open models improve, and for the emergence of more specialized agent harnesses that collect valuable training data.
Tweet Feed
AI Model Releases & Comparisons
@kunchenguid · 2026-07-29T17:56
now that the dust is settling around the last wave of model releases, and i've had enough time to use all these models in practice, let me share a more complete set of thoughts
- grok 4.5 is probably the single most significant event during the last couple of weeks
i've been talking with many heavy users across model families, and it's pretty much a consensus that grok 4.5 is the most "pleasant" frontier model to work with day to day. it's incredible how precisely the team behind it found this perfect sweet spot and created a model that's so fast, efficient and capable
it proved spacexai is now a 3rd real player in addition to anthropic and openai. they have real time data from the biggest public townsquare of humans, they have acquired a popular agent harness, and now they have proven they can build great models. and if you look closely, they are designing their own chips, they have their own data centers, they can send GPUs into the space and create tokens out of sunshine
holy shit
- speaking of data, human usage over a harness proved to be extremely important for training good models. grok 4.5 was the first model that incorporated cursor's data and it made a massive difference compared to previous generations of grok
this explained why amazon mandates employee usage of kiro, why meta installed mass surveillance over employee devices, why anthropic bans 3rd party harnesses, and why google is still struggling with gemini - because they don't have a popular harness with mass adoption to collect the data
this is part of why i don't think the subsidized LLM subscriptions will end any time soon, because a wide consumer adoption is the best source of data collection. we're paying the subsidized tokens by teaching their models how to get work done
- opus 5 flopped. almost no one likes it. the only people who like it seem to be using it to one-shot 3d games that look impressive but no one will ever buy
anything that AI can one-shot is just the definition of garbage, because if you can one-shot this thing with a quick prompt, you should know that it means billions of other people can also do it - you will not create anything of value this way
and this is just a symptom of a more fundamental problem that model training is heading down a slippery slope where machine verifiable outcome is dominating over human feedback
the latest training process rewards the agent for running for a long time and finishing a complex project, yet no longer seems to care about how the agent talks to its human
jargons, walls of text, "an honest mistake" - opus 5 showed us that we need AI that's more human friendly. let's not build a world where we end up working with robotic a**holes all day
- fable 5 remains undefeated as the upperbound
i talked about this in my previous post about wisdom vs diligence. the benchmarks blend both together so it's not easy to see, but fable 5 is the GOAT on the "wisdom" dimension despite it not winning on every benchmark. if you used it meaningfully, you know what i'm talking about
that said, it seems anthropic is extremely paranoid about other players, including open models, reaching the same level of intelligence, which is an indication that the moat is not strong. kimi k3 is just a preview of what it looks like
at the same time, fable is the first time token cost is becoming a very real problem. it's the only model so far that i can't afford to keep using all day. i suspect this will remain true for a while, that we have to pick and choose what tasks to give to fable-tier models, not using them as a daily driver
- openai is in an interesting position
gpt models have been great at efficiency, but now grok is also very competitive. gpt also haven't quite reached the same wisdom upperbound where fable is yet - although gpt 6 may change that
i think there are two angles for openai to pursue:
-
continue to bet on efficiency, and go after enterprise adoption while claude is too expensive and grok has a brand tax to pay there. this is a very viable strategy
-
or.. compete head on with fable on wisdom and win against them on the "human friendly" aspect. although traditionally this hasn't been the strength of openai models either, so this feels like a low-ROI option
alright, that's a bit of a long post but the landscape is just becoming increasingly complex. hope these thoughts are helpful in terms of providing a reference for how to rationalize everything happening → tweet link
@jezell · 2026-07-29T18:32
RT @FT: FT exclusive: Google DeepMind has dismantled the team behind AlphaFold, its Nobel Prize-winning AI system for predicting protein st… → tweet link
@ivanfioravanti · 2026-07-29T16:00
A new Video model from @MiniMax_AI and @Hailuo_AI is here! Let me see if I can try it 😎 → tweet link
@ivanfioravanti · 2026-07-29T16:16
Only 8 x DGX Sparks 👀😱 → tweet link
@ivanfioravanti · 2026-07-29T15:52
GLM Coding Plan is supported! Let's get ready for GLM 5.5 in Pi. 😉 → tweet link
@ivanfioravanti · 2026-07-29T15:44
RT @UnslothAI: Kimi K3 can now be run locally! ✨
The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% siz… → tweet link
@Ex0byt · 2026-07-29T15:30
sebastian breaks down Kimi K3 as just a massively scaled iteration of last year's Kimi Linear architecture → tweet link
@victormustar · 2026-07-29T08:35
RT @_akhaliq: A.X K2 just dropped on Hugging Face
Large-Scale Sparse MoE (688B / 33B Active)
https://t.co/nVr17DscD3 https://t.co/zOHc6in… → tweet link
@ivanfioravanti · 2026-07-29T12:51
A 2-bit quantization delivering ~100% of FP8 on a MoE model like Qwen 3.6-35B-A3B??? 👀
I need to try this on my 3090, as soon as @luceboxai will deliver it to me 😂 → tweet link
@gdb · 2026-07-29T03:36
5.6 for solving another longstanding open problem, this time in probability: → tweet link
@badlogicgames · 2026-07-28T22:27
been using fable for a somewhat large design.
2k loc markdown file. another 2k loc of sources as context.
it - makes shit up instead of reading sources - modifies the design without the change ever having been discussed kr approved - suggested to walk individual rows in a table, one query per row - replies only in "i'm very smart" mode with the most insufferable linkedin voice - falls apart after > 200k in the context window
blew $500 on this. absolutely not fucking worth it. what have they done to our boi? → tweet link
@ivanfioravanti · 2026-07-28T20:19
Look at what Opus 5 + GPT 5.6 Sol can do! 😱 → tweet link
Developer Tools & Open Source
@gdb · 2026-07-28T22:41
we've just open-sourced the Codex Security CLI: https://t.co/gIm9X2wDdh → tweet link
@jezell · 2026-07-29T05:05
RT @OpenAIDevs: We're introducing two new transcription models in the API:
• GPT-Live-Transcribe: built for low-latency live transcription… → tweet link
@sqs · 2026-07-29T09:30
I switched @AmpCode's dictation from gpt-4o-transcribe to the new @OpenAI gpt-transcribe model on the livestream - and built a new waveform suggested by a listener https://t.co/ql7pwlm5mN → tweet link
@jack · 2026-07-29T17:36
RT @hot_town: My favorite part about Buzz?
Anthropic and OpenAI won't be happy with it.
You can change which harness + model powers your… → tweet link
@jack · 2026-07-28T20:23
RT @Hello_World: Buzz Desktop v0.5.0 is out! 🐝
Invite people with use-limited links, filter search with from:/in:/before:, bring your own… → tweet link
@jack · 2026-07-28T21:15
RT @carlosmarcialt: Buzz added Hermes Agent support today. Claudio, my main AI assistant, is already inside.
Claudio runs on my VPS with h… → tweet link
@jezell · 2026-07-29T17:44
Kata Containers 4.0.0 Brings You a New Rust Runtime https://t.co/SCNmcCeLkF → tweet link
@jezell · 2026-07-29T17:23
RT @mitchellh: I've started a new company: @superlogical! We're going to begin by building a terminal multiplexer. The entire vision is muc… → tweet link
@badlogicgames · 2026-07-29T11:46
RT @dsp_: Thrilled to announce the MCP 2026-07-28 release. It is one of MCP's biggest yet, built on 18 months of lessons learned.
A few hi… → tweet link
@badlogicgames · 2026-07-29T11:47
RT @bentlegen: 🛠️ hunk 0.17.0 is out
Patches from 10 new contributors, w/ small fixes & quality-of-life improvements like:
- official hom… → tweet link
@louszbd · 2026-07-29T16:05
It's been great to hear from so many of you that ZCode feels good to use. Here are some updates. You can now: - Open and use web pages without leaving the app. - Preview PDFs and zoom in or out, better stay in flow when reading papers. - Open large files more reliably. - Use repo wiki to generate a catalog of architecture, with diagrams added if needed. v3.5.3 improves cache hit rate. Interactions are more responsive. Enjoy! → tweet link
@Teknium · 2026-07-29T16:55
Hey Hermes - Replace your Google Home, Alexa, whatever you got with Hermes Agent on any device that can run the GUI or CLI locally!
Voice activation is now in Hermes Agent along with many other voice chat improvements.
To toggle listening for wake word on, just click the ear icon in the prompt box in your GUI, and you can activate Hermes voice chats with "Hey Hermes" - and it can even pick up profile names to activate voice chats with them! → tweet link
@Teknium · 2026-07-29T05:34
Voice chats in Hermes Agent are now way faster, most API provided TTS backends will stream their speech so you get closer to real time interactions! https://t.co/2xGZ2cZY4W → tweet link
@jxnlco · 2026-07-28T21:49
SKI GAME PROMPT
Build me a complete, polished browser game called ALPINE RUSH in ONE single index.html file. It is an endless downhill snowboarding runner with stylized low-poly alpine visuals, AAA-feel post-processing, and fully procedural audio. No build step, no frameworks, no asset files. Use three.js 0.160.1 loaded from the jsdelivr CDN via an import map... → tweet link
@MengTo · 2026-07-29T00:44
I keep seeing the same pattern with Three.js games, shaders, p5.js, and now 3D UI.
Give a coding agent a strong reference, the live URL, and the source repo. It can study the mechanics, recreate them, and give you a working place to experiment.
For this one, I turned a Trevor Noah-inspired book interaction into fictional manuals for Codex, Claude Code, and Cursor. → tweet link
@MengTo · 2026-07-29T14:52
I love the rise of free UI libraries built by designers who really care about the details.
Beautiful UI has 17 polished patterns for AI-native products, from thinking and streaming states to tool calls, approval cards, and diffs. Every component includes the source code, so you can give it to your agent and adapt the interaction to your own product. → tweet link
Local AI & Hardware
@TheAhmadOsman · 2026-07-28T22:24
We built a Local AI Hardware Arena using ODS
LLM races between
- RTX PRO 6000
- DGX Spark
- Strix Halo
- M5 MacBook Pro
- The Cloud (ChatGPT)
Let us know what you wanna see next
We will make Local AI The Default https://t.co/o3qHeNFrLb → tweet link
@TheAhmadOsman · 2026-07-29T15:39
ODS v2.6.0 is now out
Our goal with every new release is to make setting up Local AI so seamless and easy that t becomes the default
ODS is the missing layer that is much needed to make Local AI The Default for everyone out there
Let us know what you think! https://t.co/nsS0z3sW6B → tweet link
@sudoingX · 2026-07-29T00:50
CALLED IT! laguna s 2.1 is the new king of the single dgx spark for agentic coding, and it wasn't a decision but a brutal knockout.
i put both on one dgx spark, same nvfp4 with vllm engine, and ran them head to head. laguna loads to 67 gigs and just sits there comfortable, 128k context, spec-decode running, fifty gigs of headroom to spare.
step 3.7 flash is a 198b model, 113 gigs of weights, and when i went to serve it the weights loaded for thirteen minutes and then the whole spark crashed... → tweet link
@sudoingX · 2026-07-29T17:20
a few hours ago i crowned laguna s 2.1 the king of the single dgx spark, and it earned the belt, but let's be honest about how.
stepfun 3.7 flash never even turned on. it loaded weights for thirteen minutes then pulled the whole box down, twice, never served a token...
so today laguna gets a real opponent. qwen 3.5 122b, the a10b active variant, a few months old now, in nvfp4... → tweet link
@sudoingX · 2026-07-29T01:18
i've never wanted the second dgx spark more than i do right now man. the single spark showed me its ceiling this week, over and over.
i'm done benchmarking around it. i want to go through it. i keep hitting the same wall, and this morning it got personal.
i tried to load a model onto my spark that just wouldn't fit, watched it crawl for thirteen minutes and pull the whole box down, twice... → tweet link
@sudoingX · 2026-07-29T01:08
i'll say it plainly, hermes agent desktop is the best agentic app i've used, and i'm a little mad i didn't find it sooner.
it auto see the models i'm serving, laguna s 2.1 sitting on my dgx spark and the desktop app just picks it up over the tailnet and streams inference straight into the gui, no config, i pick the effort from minimal to ultra and go...
@NousResearch keeps quietly building the thing everyone else is bolting together out of five tools. this is genuinely the one. → tweet link
@sudoingX · 2026-07-28T21:51
this one nobody will tell you: run your own private git server and make it your agents' memory layer.
the idea is simple every agent on every machine you own reads and writes its memory, its rules, and its state to the same git repo, hosted on a box you control.
an agent learns something or changes a file, it commits it. bigger changes go through a pr. now every agent, on every node, shares one versioned brain... → tweet link
@ivanfioravanti · 2026-07-29T05:54
Apple Containers 1.2 released a few hours ago. It's now really stable. I still hope GPU access will be enabled in a way or another.
https://t.co/Dy3yWGiyeC → tweet link
@FrameworkPuter · 2026-07-29T18:18
More Framework Laptop 13 Pro reviews are going live, including @notebookcheck with an Editors Choice Award. https://t.co/UELCmRUvot → tweet link
@FrameworkPuter · 2026-07-28T19:55
It turns out you can't upload videos longer than 12 hours to YouTube, so our Framework Laptop 13 Pro full length battery life videos have to be stop motion speed-ups instead. We have the first one uploaded, showing >20 hours with Netflix 4k streaming in Windows. https://t.co/vEL0tx5I1s → tweet link
@LinusEkenstam · 2026-07-29T17:51
Current domestic robot staff
Excluding my production robots (3D printers) and security robots (AI powered security cameras).
I think this landscape will completely explode.
Multiple cleaning droids at home will be standard, self cleaning and mopping.
Lawn and perimeter control robots for property.
All types of maintenance robotics, window cleaning, pool keeping, wall cleaning, garage cleaning.
After that, home humanoid robots will be added and cover every other generic task. → tweet link
@LinusEkenstam · 2026-07-29T13:54
I have +4 physical robots working for me. That number is going to grow to 6 or 8 before end of year.
These are single purpose robots, but the next big breakthrough will be general purpose robotics.
Robots that can do more general task. At first there will be a limited library of tasks, but this will accelerate so fast that within very little time, general robots will be able to do almost any task a human can do.
2027-2030 will see more change than what we've seen since the birth of the new millennium. → tweet link
@LinusEkenstam · 2026-07-29T17:56
zero-delay auto-aim system on an $8 microcontroller
time to put this on my Robot Lawnmower and attach a very strong laser. https://t.co/2J7OxGdbqv → tweet link
AI Industry Analysis & Agent Workflows
@kunchenguid · 2026-07-29T04:06
i always disable all permission checks for my agents - i don't even do auto review. many people saw that in my videos and asked me how i dare do that, and that's actually a really good question
on a high level, there are 3 things that allowed me to do that -
- a mindset shift - my machine is not mine
- fully reproducible config
- isolate all secrets → tweet link
@swyx · 2026-07-28T20:20
re: hiring right now
it's a huge bull market for AI-native IC's/player-coaches it's a huge bear market for "heads of X" managers
never seen such furious bifurcation. to oversimplify: 1 year experience managing 10 agents > 10 years experience managing 10-100 people → tweet link
@TheAhmadOsman · 2026-07-28T20:15
Gatekeeping by "frontier AI companies" dressed as an innocent call for OTHERS to slow down
Regulatory capture via a new narrative now that the fear mongering one failed
DO NOT fall for this https://t.co/9IKi9A09Me → tweet link
@TheAhmadOsman · 2026-07-29T10:18
I am still in awe of being able to self-host frontier intelligence like Kimi K3 myself
Closed frontier labs crazy stupid valuations must be doing terrible now that their researchers are asking for a "slow down" before they IPO lmao → tweet link
@sudoingX · 2026-07-28T19:59
i walk through malls and cafes and watch people work, and almost every open laptop has a spreadsheet on it. someone hand formatting cells, dragging the same formula down, coloring rows one at a time, doing everything by hand what an agent would finish in a sentence.
and it hits me every time how small the x ai bubble we're in really is... that's not a knock on them, it's the opportunity. the ai bubble on this app is a rounding error next to the number of people still coloring spreadsheets by hand. the wave isn't us. it's them, the day they realize the machine already on their desk can just do the boring half of the job. → tweet link
@badlogicgames · 2026-07-29T16:21
recommended reading. being in the business of shipping software for artists, it's a pretty good analogy.
it's also why i think the "generate single digital artifact based on millions of examples" (music, images, videos) and "coding agent trained on all the intermediary traces up to the final artifact plus RL" might be what we'll be stuck with for quite a while.
there is simply not enough material to train on in other fields to get a capable model.
which is why all the big labs have turned to "everything is coding agent shaped". that works for some tasks outside software, but not enough to replace entire departments of humans with machines. → tweet link
@badlogicgames · 2026-07-29T15:28
still finding tons of value in human feedback for e.g. system designs.
agents still seem to suck keeping the whole thing in their tensors. → tweet link
@thdxr · 2026-07-28T21:46
we spend a lot of energy and money helping opensource models get better
but for coding there are still large gaps between them and the frontier
maybe raw intelligence is the comparable but zooming into day to day usage long way to go before they get the same level of polish → tweet link
@Teknium · 2026-07-29T05:10
Open models bring a lot to the table when closed models close more than just the weights → tweet link
@levelsio · 2026-07-28T21:12
More interesting things happening in how AI is radically transforming things this month but now in tech hiring, very interesting → tweet link
@levelsio · 2026-07-29T14:45
Just today I've already seen Wispr Flow, Granola and WHOOP all "reverse engineered" and open sourced with a fully free version
Very interesting to see what's happening
The question is if normies will pick up on this (I think they will) and how companies will react and pivot to still make money → tweet link
Research & Academic Access
@gdb · 2026-07-29T17:46
Science moves faster when more researchers have access to the best tools.
Getting frontier AI into the hands of academic researchers means more shots on goal against humanity's hardest problems — and more breakthroughs that benefit all of us.
Excited to see what they discover. → tweet link
@jsuarez · 2026-07-29T00:21
RT @mcbeukman: 1/ We are happy to announce the largest known dataset of expert trajectories for physics-based tasks, containing over 11M un… → tweet link
@victormustar · 2026-07-29T16:29
RT @askalphaxiv: alphaXiv 🤝 Hugging Face
We're excited to share that you can now sign in to alphaXiv with your HuggingFace account! 🤗
Tod… → tweet link