← Tech / AI / IT Monitor Index Tech / AI Generated 2026-05-31 19:30 UTC

Tech / AI / IT Monitor

May 31, 2026 · Based on tweets from the last 24 hours · 137 tweets analyzed · model: ollama-cloud/glm-5.1:cloud

Daily Intelligence Briefing — Tech / AI / IT Monitor

Date: 2026-05-31 · Reporting Window: Last 24 hours


Executive Summary

OpenAI officially launched its Robotics division, evolving its world simulation research program into a full-stack hardware and ML effort focused on building useful physical-world robots. GPT-5.5 claimed the #1 spot on DeepSWE, a hard long-horizon coding benchmark, outscoring Claude Opus 4.8 (70% vs 58% pass@1) while being ~3x more token-efficient. On the open-source front, StepFun's 198B Step 3.7 Flash vision model was demonstrated running fully locally on NVIDIA's DGX Spark, and the agent framework space heated up with significant updates to both Hermes Agent (token savings, lazy tool loading, Kanban /goal support) and OpenClaw (faster cold/warm turns, smaller installs, policy conformance). The competitive dynamics between closed and open weight models continued to narrow, with observers noting the open-source lag has shrunk to approximately four months.


Key Events


Analysis

Patterns: - Agent framework competition intensifying. Hermes Agent and OpenClaw are releasing rapid-fire updates on overlapping schedules, each carving distinct philosophies: Hermes prioritizes battery-included out-of-box functionality, OpenClaw emphasizes modularity and minimalism. This divergence will likely consolidate as the market matures. - Codex as a force multiplier. Multiple independent reports of Codex handling large-scale, long-running autonomous tasks (24-hour codeports, ad-hoc codemods) suggest it's transitioning from a coding assistant to an autonomous software engineering agent. The pro plan pricing is being recognized as heavily subsidized relative to API costs. - Local inference crossing new thresholds. The ability to run a 198B MoE model on desk-side hardware (DGX Spark) with functional reasoning capabilities marks a hardware-software co-optimization milestone. The gap between what requires cloud and what runs locally is narrowing for MoE architectures. - Open vs. closed narrative shift. Multiple voices are pushing back on claims that open-weight models are falling further behind, citing empirical evidence of a shrinking ~4-month lag.

What to watch next: - OpenAI Robotics hiring velocity and first hardware teasers — this signals a major strategic expansion beyond software. - GPT-5.5 benchmark claims being independently replicated, especially the token-efficiency narrative. - Hermes Agent vs. OpenClaw feature convergence and how Codex's integrated app experience disrupts both. - Apple Metal compiler bug resolution and its downstream impact on MLX-based inference stacks. - StepFun 3.7 Flash community adoption patterns given the practical DGX Spark deployment details now public.


Tweet Feed

AI Model Releases & Benchmarks

@sama · 2026-05-31T16:07

OpenAI Robotics is hiring, looking for exceptional full-stack hardware, ops, systems, and ML engineers to help us program and manufacture robots that are useful for society. [...] Our world simulation research program, led by Aditya Ramesh (@model_mechanic), has evolved over the past year into OpenAI Robotics. → tweet link

@gdb · 2026-05-31T18:01

OpenAI Robotics is making rapid progress towards building AI that can help people in the physical world. Apply now to join the team: → tweet link

@RydMike · 2026-05-31T14:34 (RT @reach_vb)

GPT-5.5 is #1 on DeepSWE, a hard long-horizon coding benchmark 🔥 70% pass@1 vs 58% for Claude Opus 4.8. And GPT-5.5 gets th… → tweet link

@gdb · 2026-05-31T05:22

GPT Realtime 2 unlocks some real magic: → tweet link

@sudoingX · 2026-05-30T20:22

i have to show you this. a 198B stepfun 3.7 flash vision model, reasoning out loud, on a dgx spark box locally. watch it think. [...] 198B MoE, 11B active, q4 KV cache, the full 256k context, on a single 128GB spark. → tweet link

@sudoingX · 2026-05-30T19:35

i am running stepfun's new step 3.7 flash on a dgx spark right now. 198b vision model, on a box that sits on a desk. here's how to save yourself about 3 hours of head scratching getting it loaded [...] that's the 3 hours. exact working flags and the weights in the reply. → tweet link

@TheAhmadOsman · 2026-05-30T23:18

Open weight models have lagged the state of the art closed models by four months. The lag is shrinking not expanding unlike what paid influencers would like you to believe → tweet link

@TrungTPhan · 2026-05-31T14:45

Google really the new Bell Labs, using that monopoly search money printer for new inventions: ▫️Waymo ▫️AlphaFold ▫️Transformer (Gen AI) paper ▫️Willow quantum Computing chips ▫️Debug project to eradicate mosquitos ▫️Triple unskippable back-to-back-back pre-roll YouTube ads → tweet link

Agent Frameworks: Hermes Agent

@Teknium · 2026-05-31T16:53

Welcome to the Hermes Agent crew! → tweet link

@Teknium · 2026-05-31T04:19

Just want to make this clear: We didn't make Hermes Agent to be a "starts with nothing, you work it all out" agent. [...] We want Hermes to work out of the box for most people. [...] Run hermes skills config or hermes tools to disable whatever you want. → tweet link

@Teknium · 2026-05-30T21:43

Found a way to save everyone 14% on input tokens on average during read file operations in Hermes Agent! This is now on main. hermes update to access now. → tweet link

@Teknium · 2026-05-31T08:18

You can now make your Kanban jobs work with /goal. Just a hermes update away → tweet link

@Teknium · 2026-05-30T20:09 (RT @IBuzovskyi)

HERMES AGENT JUST LEARNED TO LOAD TOOLS ONLY WHEN IT NEEDS THEM. if you attach 15+ MCP servers, their schemas eat your con… → tweet link

@sudoingX · 2026-05-31T13:33

if he had eyes for engineering instead of engagement, here are a few features he could've listed instead of the skills he mocked. [...] eleven per-model tool-call parsers [...] context auto-detection [...] a seven-step repair chain [...] that's the unglamorous backend that makes local models actually usable. → tweet link

@sudoingX · 2026-05-31T13:14

the $8 closed wrapper guy is mad that a free, open source agent ships with too many skills. [...] in hermes agent i read the skill, switch it off, or fix it. in t3 i can't see a single line. which one's actually built for the user? → tweet link

Agent Frameworks: OpenClaw

@steipete · 2026-05-31T13:10

The idea of OpenClaw is always that it should be yours. It's modular and lean, only add what you need. Fewer skills, fewer tools = your agent can work more efficiently. → tweet link

@steipete · 2026-05-30T22:29 (RT @openclaw)

OpenClaw 2026.5.28 vs .27, measured: ⚡ Cold turns 14.5% faster 🔥 Warm turns 16.0% faster 📦 Fresh install 52.8% smaller 🧩 Pac… → tweet link

@steipete · 2026-05-31T07:57 (RT @joshavant)

Now in @openclaw: Use a guardian agent to evaluate the safety of your agent's proposed system calls, only prompting you when… → tweet link

@steipete · 2026-05-31T13:08 (RT @vincent_koc)

From YOLO to Auto. LLM-based auto security approvals to make things safer for our users, and works with ANY model. 🦞 → tweet link

@steipete · 2026-05-31T04:49 (RT @OmarShahine)

If you run @openclaw you should use the new policy conformance plugin - Verifiable proof things don't ever drift, a presen… → tweet link

Codex & AI Coding Tools

@jezell · 2026-05-31T17:46

Switching everyone back back to a codex personal pro plan. This 11.2B tokens was $200 / mo. [...] The pro plans are just too good to pass up when you get like $8k/week for $200. → tweet link

@jezell · 2026-05-30T22:51

Codex has been working nonstop on porting the LibreOffice object model to dart for 24 hours so far with /goal. Up to 84k loc with tests + classes. → tweet link

@steipete · 2026-05-31T15:59

Haven't seen codex writing ad-hoc codemods before, but it just did for a bigger TypeScript migration. Impressed. → tweet link

@gdb · 2026-05-31T06:54

codex computer use is viscerally compelling → tweet link

@nummanali · 2026-05-31T15:38

He is using openai-codex as the provider which means he's using his ChatGPT account in Pi. This allows for heavily subsidised tokens think 10-20x cheaper than API rates [...] OpenAIs decision to allow 3rd party usage of GPT inference has led to incredible OSS projects → tweet link

@nummanali · 2026-05-30T21:42

The Codex app is going to supersede OpenClaw and Hermes. I simply do not see the value in them when Codex can do everything in a cleaner slicker UI. The ChatGPT app enables remote app to Codex on your machine. → tweet link

@nummanali · 2026-05-30T19:08

Codex app is better than Claude Desktop for these two simple reasons: 1. Sessions are synced with terminal 2. No confusion of Chat/Cowork/Code. Why Claude Desktop hasn't been unified is anyone's guess → tweet link

@TheAhmadOsman · 2026-05-30T19:14

PRO TIP: For Codex Cli & other agents like Claude Code, Droid, OpenCode, etc. There's a crucial recipe: 1. Modularity 2. Domain-Driven Design 3. Painfully explicit specs 4. Excessive documentation. This is systems engineering. → tweet link

@thdxr · 2026-05-31T15:46

someone connect me with pewdiepie so i can get him a preview of opencode 2.0 which is better for his use case → tweet link

Local Inference, Hardware & Systems

@badlogicgames · 2026-05-31T08:56

I did my part! Created a minimal repro boiling it down to quantized matmul. Looks like a bug in Apple's latest Metal compiler. Fun! → tweet link

@badlogicgames · 2026-05-30T23:54

i can hear coil wine on my M5 MAX when i run local inference 😭 → tweet link

@badlogicgames · 2026-05-31T15:31

poor man's /goal, to port the MLX based qwen3-tts inference engine over to GGLM and C for some cross-platform goodness. → tweet link

@badlogicgames · 2026-05-30T23:27

gpt found a bug in mlx-c 0.31.2. now i wonder if i should send a slop PR, because i haven't looked into the issue myself (analysis sounds legit tho :p) → tweet link

@sudoingX · 2026-05-31T06:03

has anyone tried blender automation with smaller local models like qwen 27B or anything running on dgx spark? curious what's possible. → tweet link

@gospaceport · 2026-05-31T15:06

Jank factor ~3/5 on my Z440 transplant → tweet link

Developer Infrastructure & Tooling

@thdxr · 2026-05-31T18:11

alright all of you that maintain a cli oauth flow. i hope it's obvious to you now doing the whole browser link callback to localhost thing is dumb and annoying af in ssh. please implement the code flow that polls - try gh cli login flow to see it → tweet link

@badlogicgames · 2026-05-31T15:49

10 years ago, i built myself my own reddit web reader called ledit. [...] now they blocked raw requests without auth cookies. so i had pi write a proxy that uses apple script to spawn chrome every 30 minutes to refresh the cookies, extract them from the chrome profile, and use them to request the JSON. i will continue to win this fight. → tweet link

@LinusEkenstam · 2026-05-30T19:23

When you use Opus 4.8 to orchestrate all your open-source sub-agents… → tweet link

@hnasr · 2026-05-31T14:01

The scars behind a piece of software innovation are not transferable. The final product though, be it a data structure, an algorithm or even a protocol is packaged, labeled, given a fancy name [...] The engineer who reads about the innovation is stripped from the joy (and pain) of the exploration. → tweet link

@Teknium · 2026-05-30T19:55 (RT @nvidia)

A new era of PC. 25.0528, 121.5990 → tweet link

@badlogicgames · 2026-05-30T23:18 (RT @SIGKITTEN)

personal update: I'm joining @OpenAI to work on Codex! → tweet link

@TheAhmadOsman · 2026-05-31T03:50

Jensen tried to stop this from happening btw. Stopping NVIDIA from selling to China will go down the in history as a massive fumble in the chip supremacy war → tweet link