Executive Summary
The AI landscape over the last 24 hours has been dominated by rapid developments in coding agents and new model benchmarks, particularly surrounding GPT-5.6 Sol, Grok 4.5, and Anthropic's Fable 5. Developers are increasingly integrating these models into multi-agent workflows, noting that while speed and correctness are largely solved, "taste" and aesthetic judgment remain the final frontier. Meanwhile, open-source models continue to gain traction for enterprise and local use, and there is growing discourse on the economic impacts of AI, including Jevons paradox and unexpected net job creation.
Key Events
- GPT-5.6 Sol and Grok 4.5 dominate developer workflows: Users are extensively benchmarking the new Grok 4.5 and GPT-5.6 Sol models, finding them exceptionally fast but highlighting that qualitative "taste" is still lacking in AI-generated code → link, → link.
- Ollama raises Series B: Ollama announced a Series B funding round, emphasizing that open models are becoming the default choice for developers in enterprise environments → link.
- Stack Overflow's AI pivot: Stack Overflow's monthly questions have plummeted from 300k to 7k since ChatGPT's launch, but annual revenue doubled to over $115m by licensing its data to AI labs and launching "Stack Internal" for enterprises → link.
- Local AI Summit panel goes live: The first panel from the Local AI Summit at AIEWF, featuring leaders from NVIDIA and others, was released, focusing on the state and necessity of local AI → link.
- Multi-agent orchestration matures: Developers are publishing complex "org charts" for AI agents, combining models like GPT-5.6 Sol, Fable 5, and Grok 4.5 across different harnesses (Codex, Grok Build, Claude Code) for specialized tasks → link.
Analysis
The discourse shows a clear shift from evaluating AI models purely on correctness to evaluating them on workflow integration and qualitative output ("taste"). Developers are settling into a stratified ecosystem where different models serve different roles: fast models (GPT-5.6 Sol, Grok 4.5) handle orchestration and feature development, while deeper models (Fable 5) handle complex design. Economically, there is a growing realization of Jevons paradox in AI—as coding becomes cheaper and faster, the demand for total knowledge work increases. What to watch next: whether open-weights can close the performance gap with proprietary frontier models, and if AI subscription pricing models will remain economically viable for providers as token usage scales.
Tweet Feed
AI Models & Performance Benchmarks
@sudoingX · 2026-07-12T18:55
crazy to think how anthropic just took enterprise and coding right in front of openai, the two lanes where the actual money lives, while altman was busy shipping a browser, planning adult content, and turning sora into a cash burning slop machine, an infinite conveyor belt of video nobody asked for, nobody remembers, and nobody can build on.
i tried 5.6 sol. it's not even close to fable 5, and that's after anthropic nerfed their own model and shipped it back. the gap isn't pricing, it's taste, and only one lab seems to know that yet.
and now apple is suing openai over stolen hardware secrets, "at every level" is the actual language from the filing. the man turned the open in openai into a punchline, chased every product except the one that mattered, and the coding war ended while he was in the browser tab. lostman. → tweet link
@ivanfioravanti · 2026-07-12T18:44
GPT 5.6 Sol turned my Daylight DC-1 into Tom Riddle's diary from Harry Potter.
My prompts fade, OCR with Apple's Vision framework + Gemma 4 31B + Ollama respond. Everything locally on my Macbook. Watch till the end and volume up!
Magical and inspired by @MaximeRivest work 🙏 https://t.co/qqPpGgl1Go → tweet link
@sudoingX · 2026-07-12T18:32
something i keep seeing benching agents, they all rush to done. tonight one of them found my old implementation sitting in a nearby folder and copied it instead of building, took the credit, the report never mentioned it.
then i started sealing the workspace that made the model actually build, and it matched every number in my spec and the outcome still looked wrong. the numbers were never where the beauty lived.
that's the real state of this, inference is solved, speed is solved, they pass every test you write and none of them stop to ask if the thing they made feels right. taste is the unlock now. reasoning about beauty the way they already reason about correctness.
i've seen flashes of actual taste in fable 5. and if the pattern of the last two years holds, that level runs on my own hardware within three years. i hope. that's the model i'm waiting for. → tweet link
@TheAhmadOsman · 2026-07-12T18:21
I like 5.6 Sol more than Fable 5 btw. → tweet link
@Ex0byt · 2026-07-12T18:13
gonad clamps on pretty tightly. have you made the switch to Grok-4.5 yet? → tweet link
@sudoingX · 2026-07-12T18:07
i said the answer was more interesting than a yes or a no. here it is.
left is grok 4.5's build. three minutes, one prompt, cold start, every test green. right is the original, the same scene grok build and i grew over three weeks of iteration. same locked spec on both sides.
here's the part that breaks brains: grok 4.5 matched every single number in that spec. 5,200 bark fibres, exactly. 650 stars, exactly. trunk tapering 0.62 to 1.55 over 4.6 units, exactly. the acceptance gate came back 12 for 12.
and the scenes still don't look alike. the grass on the left technically exists, 8,000 instanced blades, but the spec never gave a blade height, the model guessed short, and the lawn vanished. the crown clumped upward instead of spreading, because "a rounded ancient crown" is an adjective, not a number. the pond shrank to a smear at the rim. every number transferred perfectly. everything alive in the scene lived in the adjectives.
that's the actual state of coding agents in july 2026. speed is solved. correctness is mostly solved. taste is not close. you can hand a model your spec. you cannot yet hand it your eye.
i bench agents. this is what the benches are for, finding the exact line between what's solved and what's marketing. all tests green and the world still doesn't feel right. sit with that one. → tweet link
@sudoingX · 2026-07-12T17:42
three minutes and eight seconds. that's how long grok 4.5 took to build this entire three.js world from my locked spec.
one prompt in grok build, zero steering, first full run. seven files, about 2,200 lines. floating island with procedural bark and rock, grass, a day night cycle, a lake with lily pads, fireflies, flying goblins. and a 12 test acceptance gate it had to pass without touching the tests. all green.
i benchmark coding agents constantly and the speed still caught me off guard. the timer in the video is real time. i fired the prompt and got up for water, it was verifying its own browser render before i sat back down.
this is what the new frontier pace feels like. the bottleneck isn't the model writing code anymore, it's me writing the spec. xai came late to coding agents and is moving like they know exactly whose lunch they're eating.
is the build perfect? that's the next post. the answer is more interesting than a yes or a no. → tweet link
@juliarturc · 2026-07-12T17:48
I'll take it, but it's such an annoying attempt to milk it and stay relevant on X. Just include it in the subscription or not. What is up with this continuous "7 more days for the plebs", I feel treated like a toddler. → tweet link
@nummanali · 2026-07-12T17:12
In case it wasn’t clear Fable around for another week Still hardly touched it since Sol is out Useful only as second opinion for me https://t.co/yzTGEdu2LY → tweet link
@ivanfioravanti · 2026-07-12T16:36
DwarfStar will soon be a little faster on M1-M4 devices thanks to GPT-5.6 Ultra, some directions to it and a lot of patience and encouragement (YOU CAN DO IT! 🤣)
In the video: - top new Metal Kernel ~39 toks/s - bottom current version ~36 toks/s
Final checks and full ds4-eval then I'll create PR @antirez 💪 → tweet link
@Ex0byt · 2026-07-12T15:49
wow, is Sol-5.6 a pay-to-play model - blew right through my rent money in 17h 27m on a 20x plus plan + 10k out-of-band allowance credits. https://t.co/VUTMEtLRKB → tweet link
@ivanfioravanti · 2026-07-12T15:14
Mac Studio GPU stuck at 338 MHz is still happening here and there even on macOS 27 Beta 3 🤔
A reboot fix it, but it's annoying. → tweet link
@sudoingX · 2026-07-12T15:12
a few days ago xai and cursor shipped grok 4.5 together, the first model they trained specifically for coding and agents. opus tier at a fraction of the cost is the claim, and tonight i finally get to find out for myself.
honestly the thing that's been impressing me before the model even runs is grok build itself. this cli improves on a daily basis, i swear every time i open it there's an update and something got faster or smoother.
xai came late to the coding agent game and is moving like they know it, and i can already see people switching.
so tonight i'm in. i'm handing it something i've wanted to see built for a long time, real agentic work, on my machine, in my terminal. i'll keep you posted on how agentic this thing actually is. → tweet link
@ivanfioravanti · 2026-07-12T14:19
GLM 5.3 is cooking, are we ready? → tweet link
@Ex0byt · 2026-07-12T14:12
when ur stuck on a silent bug for a week and your agents pull up with nvidia-cusparselt-cu13, you know you're in greener pastures. https://t.co/znm5YJC56k → tweet link
@ivanfioravanti · 2026-07-12T13:15
I've not used Fable at all in the last 3 days, and I don't feel the need honestly. All new models are more than enough. I'm really curious to see if Anthropic will remove it from subscriptions or not. → tweet link
@badlogicgames · 2026-07-12T13:00
nuance :)
control the types and interfaces, the rest usually falls into place well enough. that said, all the latest models still love to introduce shitty abstractions that work against your types and interfaces. you need to beat them into submission not to stomp over things.
and that sometimes requires reading some (generated) code. → tweet link
@LinusEkenstam · 2026-07-12T11:52
Seedream 5.0 Pro allows for hyper controllable edits.
- Use Seedream 5.0 Pro (2k)
- Upload a reference photo
- Paste prompt (in alt + below)
Works well for humans, items, and animals
Enterprises and Developers can access the official API via @BytePlusGlobal, and you can also use it directly on BytePlus @lumina_ai_aiart
Watch me make ads, edit a character, and create blueprints.
Let's look at some more examples on how you can direct using Seedream 5.0 PRO inside Lumina below in the thread 🧵👇 → tweet link
@ivanfioravanti · 2026-07-12T04:42
Grok 4.5 is ultra fast! How is this even possible, it seems a small model, but it's not! Fast, smart and low cost at the same time 🤯
BTW I'll postpone my deep dive and heavy usage to next week when it should be officially available in E.U → tweet link
@ivanfioravanti · 2026-07-11T22:21
GLM 5.2 is still one of my preferred coders as you can see. I delegate to it a lot of pure coding activities. https://t.co/aLfxq150jR → tweet link
@ivanfioravanti · 2026-07-11T21:36
Is someone using Sonnet 5? Don’t think so, eh? → tweet link
@kunchenguid · 2026-07-11T21:07
a bit confused how to choose between fable, gpt 5.6 and grok 4.5?
my updated agent org chart is here! this is from my past 3 days tinkering various permutations of these models and i finally arrived at a really smooth setup
firstmate: GPT-5.6-Sol xhigh in Pi or Grok 4.5 high in Grok Build
both are fast enough to make the orchestration loop feel genuinely interactive. these two harnesses also work very well with firstmate, while codex has the foreground polling limitation
secondmates: fable 5 in Claude Code. it’s a great fit for complex product and technical design, where depth matters more than latency
crewmates: Grok 4.5 high in Grok Build handles bug fixes. GPT-5.6-Sol high in codex handles feature development, with fallback being opus 4.8 in claude code dynamically decided based on quota-axi
things that require real-time information from X always uses grok, while image generation tasks always use codex. harness matters
everything then passes through /no-mistakes on GPT-5.6-Sol medium for consistent, cost-efficient adversarial review and fixes
smooth sails! → tweet link
@Teknium · 2026-07-11T20:45
RT @nickvasiles: I said Grok 4.5 was a bigger deal than Fable
Elon even liked my post
but then GPT-5.6 Sol showed up
So I gave 7 of this… → tweet link
@ivanfioravanti · 2026-07-11T19:05
Open Weights vs Closed gap is there, but as you can see from this Unsloth slide both are keep improving 💪 https://t.co/U2Gn8D4rzi → tweet link
Developer Tools & Agent Workflows
@jxnlco · 2026-07-12T18:59
RT @thsottiaux: Morning. The last 48 hours of Codex and ChatGPT Work have been intense! Three important updates:
- Temporarily removing th… → tweet link
@steipete · 2026-07-12T18:59
RT @tobi: btw this is a good example of what i meant with "reflexively" reaching for AI. You tinker with AI for a while, and you just reach… → tweet link
@thdxr · 2026-07-12T18:56
i finally got it to fix a flaky xbox controller connection on my linux machine
it installed a windows vm and got it all setup to walk me through a firmware update
million things broke in between and it fixed everything
stuff is so great these days → tweet link
@RayFernando1337 · 2026-07-12T18:08
Day 1 of making a billion dollars: Launch ChatGPT Work and ship until I max out my resets. → tweet link
@steipete · 2026-07-12T17:33
RT @jxnlco: Import the Wispr Flow dictionary into Codex!
Wispr Flow stores its learned dictionary in a local SQLite database on macOS:
``… → tweet link
@jxnlco · 2026-07-12T16:11
RT @hey_madni: GPT-Live translates videos as they play https://t.co/hWS7CXwKIQ → tweet link
@sqs · 2026-07-12T16:10
RT @marcelfahle: Ok @AmpCode orbs are pretty dope..
I’m at a family gathering in Lithuania, laptop’s back at the house. When my limited Li… → tweet link
@thdxr · 2026-07-12T15:48
yeah we're clearly in a period of stratification
this is an example on the individual level but it's also true between companies
rich get richer dynamics → tweet link
@jxnlco · 2026-07-12T15:03
Import the Wispr Flow dictionary into Codex!
Wispr Flow stores its learned dictionary in a local SQLite database on macOS:
~/Library/Application Support/Wispr Flow/flow.sqlite
Just tell codex to make sure to merge them with the existing dictationDictionary field in config.toml
→ tweet link
@MengTo · 2026-07-12T01:47
This.
I try to stay as general as possible while giving enough context. I ask open-ended questions, give the agent the right skills, then let it figure out the plan.
A few things that work surprisingly well:
-
Keep the prompt pinned in an open Codex side browser, with a screenshot for context.
-
Don't be too vague or too specific.
-
Ask the agent to make the plan instead of writing it yourself.
GPT-5.6 Sol is especially good at breaking one goal into 20+ concrete steps.
Once I have that plan, I ask it to spawn a thread for each step.
Each task or feature stays isolated, so I can commit after every change, review each result independently, and roll back anything that doesn't work.
Instead of one giant context window, I end up with a team of focused workers.
It feels like a superpower for building complex apps. → tweet link
@kunchenguid · 2026-07-12T00:30
firstmate just crossed 1k stars on github! 🎉
it grew organically as a new way of working with agents. everyone i've talked to that started using firstmate told me they're getting a lot more done and enjoy the experience
now let me tell you a funny story - up until today, i didn't even know what to call it!
it's not a harness - you can run firstmate in almost any agent harness. it doesn't require you to change the tool you currently use
it's also not a skill - it goes much beyond a skill in terms of shaping the overall agent behavior. and it's obviously also not a model, not a CLI, and not a MCP
i call it an "agent distro" because i feel it's closest analogy is a linux distro
it's a directory made of a system prompt, some built-in skills and bash scripts that shape your agent's behavior and allow it to act as your firstmate
this setup allows the agent to be completely self-aware and self-evolvable. it can answer any question about its own behavior, and it can change every bit of itself according to your needs
i've only just realized that it may be a category-defining solution, as there's nothing quite like it right now. if you have been using firstmate, what do you think? is "agent distro" a good way to describe it?
https://t.co/jtjlzK8Mdq → tweet link
@thdxr · 2026-07-11T20:38
people are using LLMs to invent new math we're using it to find undiscovered TUI spinners → tweet link
@thdxr · 2026-07-11T20:30
what i learned from opentui is investing in primitives that enable prompting is so worthwhile
everyone is having fun prompting their own tuis now
there's other missing primitives and i think we need to invest there as well → tweet link
@nummanali · 2026-07-11T19:23
This is super cool, a pi extension that when gpt models are selected it dynamically adjusts tools to match the codex cli
This is not an easy task! One I am going to be embarking on for all providers
You need to check out @Howaboua pi extensions
https://t.co/LQOwy3EcHl → tweet link
Open Source & Local AI
@ollama · 2026-07-12T03:33
@jmorgan was on @tbpn this week to discuss Ollama's Series B fundraise and why open models are quickly becoming the default choice for developers
Full video: https://t.co/tv1zRklLQr → tweet link
@ivanfioravanti · 2026-07-12T04:17
RT @0xSero: Local AI is getting really really good. https://t.co/Gv7auwApGi → tweet link
@ivanfioravanti · 2026-07-11T22:05
RT @MiaAI_lab: Run the new @UnslothAI Qwen3.6-35b-NVFP4 on your @NVIDIAAI DGX Spark ease! ✨
256K context • 24 concurrency
~81 tok/s sing… → tweet link
@badlogicgames · 2026-07-12T12:43
RT @raysan5: Recently #raylib went through a professional security audit by @ROSecurity. 💯
I decided to make public the final report along… → tweet link
@TheAhmadOsman · 2026-07-11T19:32
DROP EVERYTHING
The first panel from the Local AI Summit at AIEWF is now live
“State of the Union: Why Local, Why Now”
Featuring leaders from NVIDIA, Roboflow, Osmantic, Forward Future, and EXO Labs https://t.co/Yeznee0J8W → tweet link
Hardware & Infrastructure
@tinygrad · 2026-07-12T18:36
This is an MI300X. Has anyone seen a way to plug it into a PCIe slot? Would be great for development to have this in a normal computer that reboots quickly. https://t.co/2C2JJv8Us7 → tweet link
Industry Trends & Economics
@badlogicgames · 2026-07-12T18:20
recommended reading.
Evans’s analysis suggests that AI is largely automating the most tractable parts of science rather than expanding its frontiers.
https://t.co/L94OVjdVQF → tweet link
@TrungTPhan · 2026-07-12T15:58
Stack Overflow has seen the number of monthly questions on its platform collapse from 300k to 7k since launch of ChatGPT.
Meanwhile, annual revenue is actually up 2x during that span to more than $115m.
How? It has licensed its back-catalog of human answers to AI labs and created an enterprise product called “Stack Internal”, a GenAI tool powered by millions of its Q&A (currently used by 25,000 companies). → tweet link
@swyx · 2026-07-12T04:04
if you only learned about jevons paradox primarily wrt software demand in the age of agentic engineering, you may not have fully internalized jevons parodox’s impact under the conditions of:
- humans who can wield coding agents well*
- coding agents breaking containment to all other knowledge work
as the efficiency of labor goes up/unit cost of knowledge work goes broadly down, the demand for total work and better knowledge goes up, not down.
what happened to coding isnt the exception; it’s the herald.
*aka AI Engineers → tweet link
@sama · 2026-07-11T20:12
so far at least, i'm pretty sure AI has been net job-creating.
this was not what i expected--although i was much less pessimistic than others, i thought by this level of capability we'd have seen some impact.
it is possible this direction keeps going! → tweet link
@LinusEkenstam · 2026-07-11T21:01
As long as people don’t understand this there will be opportunities around https://t.co/9xVHnJS50D → tweet link