Executive Summary
The AI and tech ecosystem saw several massive developments over the last 24 hours, headlined by NVIDIA's announcement of a $12.9 billion acquisition to join forces with Hugging Face, signaling a major consolidation of power behind open-source AI. Simultaneously, developers faced widespread disruptions as APIs from OpenAI, Anthropic, and xAI all went down at the same time. Meanwhile, rumors of OpenAI's next-generation "GPT-6-Astra" models surfaced via leaked Codex endpoints, and Perplexity open-sourced Lily, its high-performance local inference engine. The day also highlighted the growing threat of sophisticated AI fraud, with a deepfaked CEO attempting to scam a cloud inference provider.
Key Events
- NVIDIA acquires Hugging Face: NVIDIA and Hugging Face announced a $12.9 billion acquisition to join forces, cementing open-source AI and open weights as the dominant market strategy. → link
- Simultaneous AI API outage: Developers reported that OpenAI, Anthropic, and xAI APIs all experienced downtime at the exact same time. → link
- OpenAI "GPT-6-Astra" endpoints leak: New model IDs for "gpt-6-astra" and "gpt-6-astra-aeon" appeared in Codex, suggesting an imminent rollout. → link
- Perplexity open-sources Lily inference engine: Perplexity released Lily, its local hybrid compute inference engine, which reportedly offers 1.35x faster inference than MLX-LM. → link
- Deepfake CEO scam attempt: A scammer used a deepfake of a CEO on a video call to try and骗 increased rate limits from an inference provider, highlighting new AI-driven social engineering threats. → link
Analysis
The tech landscape is showing rapid bifurcation between proprietary frontier models and an increasingly potent open-source/local ecosystem. The NVIDIA-Hugging Face deal is a massive validation of the "own your weights" philosophy, providing the hardware backing to scale open models globally. Meanwhile, the simultaneous outage of major proprietary APIs underscores the fragility of centralized cloud AI and acts as a catalyst for local-first AI development. We are also seeing a surge in highly capable, deeply quantized open models (like 1-bit 27B models running on 6GB VRAM) that make frontier-level intelligence accessible on consumer hardware. Watch for OpenAI's "Astra" rollout to potentially re-shift the benchmark landscape, and expect local inference engines like Lily, Ollama, and custom kernels to capture more developer mindshare as cloud reliability is questioned.
Tweet Feed
Industry M&A & Strategy
@TheAhmadOsman · 2026-09-03T18:39
Congrats to all the friends at Hugging Face and NVIDIA 🤗💚
I'm becoming more certain by the day that Opensource AI has already won
I trust NVIDIA will continue pushing open standards and supporting Hugging Face in remaining open across platforms and hardware providers https://t.co/WxxUiP7njB → tweet link
@sudoingX · 2026-09-03T14:51
the most powerful chip company on earth just paid 12 billion dollars for the home of open weights.
open source won the market. models free, weights yours, and the hardware king is all in. unstoppable is the right word. https://t.co/RMyifZFmvY → tweet link
@alexocheema · 2026-09-03T12:15
congrats @huggingface and @nvidia!
sounds like a great outcome
im super happy about these parts in particular:
“keeping the platform open, independent and compute agnostic”
and
“empowering 100 million AI builders to own their intelligence rather than rent it” → tweet link
AI Model Leaks & Outages
@kunchenguid · 2026-09-03T15:43
what on earth could take down gpt, claude and grok all AT THE SAME TIME????!!!! → tweet link
@levelsio · 2026-09-03T15:29
No way now ChatGPT is down too?!!!
Is this the AGI takeover? https://t.co/oK7ltNz4yV → tweet link
@jezell · 2026-09-03T16:46
RT @DanDr1s: 🚨 Two new GPT-6 model IDs have appeared in Codex:
• gpt-6-astra • gpt-6-astra-aeon
gpt-6-astra-aeon likely means a higher-ef… → tweet link
@jezell · 2026-09-03T14:46
Interesting error, Astra rollout in progress? https://t.co/LQ8P50RCtV → tweet link
Open Source Models & Local AI
@sudoingX · 2026-09-03T00:01
my model recommendations for every gpu class, 6gb card to a 256gb double dgx spark cluster, from my bench book:
────── 6GB ──────
bonsai 27b at 1-bit Q1_0: 20.5 tok/s, 8K context, 3.5GB of weights. a 27b thinking on a 6gb card is the wildest small vram result i own ... [truncated for brevity] → tweet link
@thdxr · 2026-09-03T17:45
meta muse spark has dethroned deepseek as the most used model of the day
first time an american model tops this list https://t.co/fszDhCRQyV → tweet link
@sudoingX · 2026-09-03T17:16
this lab is relentless man. dropped ling 3.0 flash, then tiny, now a finance build of flash plus an expert benchmark in the same drop, all open for anyone... open weights and open evals in one release. we want more of this energy. → tweet link
@LinusEkenstam · 2026-09-02T22:02
Quite a big deal.
1.35x faster inference than MLX-LM 1.23x faster at prefill
having the hybrid option to pick from inside perplexity is awesome. seeing Lily get open sourced is even better.
Anyone can leverage this, and OS can push this even further. → tweet link
@TheAhmadOsman · 2026-09-03T02:18
You don’t “run a model” - You run Kernels
The model is just a graph The Inference Engine is scheduler / optimizer / executor
But the actual work? That happens in the Kernels ... The model is the recipe The hardware is the kitchen The Kernels are the knives, pans, burners, and the chef not cutting onions with a spoon
Most people benchmark models The real ones benchmark the Kernels underneath → tweet link
Developer Tools & Agents
@Teknium · 2026-09-03T07:22
Today I've been putting Hermes Agent to the ultimate test of dramatically cleaning up the Hermes Agent repo.
One /goal and it's been working for ~15hrs now and has simplified, unified, optimized, removed, or in some way cleaned up the repo enough to shed 375,000 lines of code. → tweet link
@thdxr · 2026-09-03T07:18
i analyzed a weeks worth of data via honeycomb + planetscale mcp and opencode2's codemode
then setup a script to email 140,000 impacted users with personalized instructions for how to fix the tool they're using
all on OpenCode's muse spark free tier → tweet link
@sqs · 2026-09-03T01:45
Orb startup is seeing a higher error rate right now. Sorry for the issues many are seeing when starting or resuming orbs. https://t.co/QkoqLw9oC3 has the latest, and we expect it to be fixed soon.
You can use an Amp runner (amp --no-tui) or the Amp CLI and mention Amp threads as an immediate workaround. We will be rolling out more orb providers soon to provide redundancy. → tweet link
@jxnlco · 2026-09-03T02:38
websites can give chatgpt work + codex tools directly through webmcp.
on supported accounts, open a supported site in the desktop app’s built-in browser. the address-bar arrow shows its tools. no separate connection needed. → tweet link
Security & AI Misuse
@thdxr · 2026-09-03T15:06
we got hit with a crazy scam attempt
someone signed up for our inference and did a video call with our sales person confirming intent to spend $400,000 over the next year and asked for increased rate limits ... messaged the real CEO on instagram and yeah he had no idea what we were talking about, the scammer was deepfaking him
be careful out there, things are getting crazy → tweet link