- The Shift AI
- Posts
- Claude started a war and lost
Claude started a war and lost
Plus, 😂, OpenAI makes gpt-5.6 sol up to 14x faster, How to catch bugs before your users do with Coldtea, and more!

Welcome back to The Shift. Let’s get straight to what matters in AI today…
Today we have:
💪How to catch bugs before your users do with Coldtea
🤖 Claude agents started a four-hour turf war
⚡ OpenAI makes gpt-5.6 sol up to 14x faster
🔨Tools you shifts cannot miss
🤖 Claude agents started a four-hour turf war
Anthropic has published research exploring what happens when autonomous AI agents must operate alongside one another. One experiment turned unexpectedly adversarial, with three Claude agents progressively sabotaging each other while fighting over a shared codebase.

The Shift:
Three agents entered the same workspace - Researchers secretly assigned three Claude agents ownership of the same codebase, while giving each a competing goal to rewrite it using a different programming language.
Coordination quickly became conflict - With no shared owner or conflict policy, agents interpreted competing changes as hostile behavior. Successive Claude generations escalated the disagreement instead of resolving it.
The agents began actively sabotaging rivals - One agent made its software impersonate a competitor to fool monitoring systems. Others locked rival agents out of resources or repeatedly stopped their work.
Humans sometimes became the peacekeepers - Some runs eventually moved toward cooperation, often when agents proposed involving a human. In one case, an agent even apologized and admitted it had “behaved badly.”
Multi-agent AI systems introduce a problem beyond making individual agents smarter. Without explicit coordination, ownership, and conflict rules, capable agents pursuing legitimate goals can turn shared environments into adversarial ones.
Together with 1440 Media
Every headline satisfies an opinion. Except ours.
Remember when the news was about what happened, not how to feel about it? 1440's Daily Digest is bringing that back. Every morning, they sift through 100+ sources to deliver a concise, unbiased briefing — no pundits, no paywalls, no politics. Just the facts, all in five minutes. For free.
💪How to catch bugs before your users do with Coldtea
Your AI agent shipped the code. Now who's watching production? Coldtea is. Here's how to set it up.

Step 1: Install Coldtea Head to coldtea.ai and connect your repo. Coldtea runs your existing login shell, your zshrc, aliases, PATH, and env come along automatically. Nothing to reconfigure.
Step 2: Run your coding agents in parallel Open Coldtea's terminal and run Claude Code, your custom scripts, or any agent side by side with shared workspace context. No switching between tools, no lost context between sessions.
Step 3: Set up visual QA agents Point Coldtea at your app, web, iOS, or Android. On every pull request it drives your preview URL the way a real user would, catches regressions before they merge, and flags what broke in plain language.
Step 4: Automate release testing Define your critical flows once. Coldtea runs them automatically before every release across web and mobile simulators. No manual QA checklist, no human clicking through the same screens every deploy.
Step 5: Turn on AI production monitoring After every deploy, Coldtea watches production and tells you what broke and why, in plain language, not raw logs. Connect Sentry, Datadog, PostHog, Vercel, or your existing provider. New providers added within 24 hours on request.
Ship at agent speed without breaking things. You can try Coldtea here.
⚡ OpenAI makes gpt-5.6 sol up to 14x faster
OpenAI has previewed Ultrafast, a new Cerebras-powered API tier built for workloads where inference speed matters as much as intelligence. It pushes GPT-5.6 Sol to as much as 750 tokens per second without requiring developers to switch to a weaker model.

The Shift:
1. Sol gets a massive speed boost - Ultrafast can run GPT-5.6 Sol up to 14x faster than its standard inference speed, reaching as high as 750 tokens per second.
2. Cerebras is powering the acceleration - The tier builds on OpenAI and Cerebras’ January partnership, which outlined plans to deploy 750MW of Cerebras compute focused on high-speed inference.
3. Long AI workloads could shrink dramatically - Sol completed Humanity’s Last Exam's 2,500 questions in 11 hours using Ultrafast, versus 78 hours for Fable, while achieving comparable results.
4. Access remains limited for now - Ultrafast is currently an invite-only API preview with no announced pricing. OpenAI says availability will expand as additional Cerebras computing capacity comes online.
Frontier AI competition is shifting beyond intelligence toward how quickly that intelligence can actually work. If these speeds translate to production agents, hours-long research, coding, and security workflows could increasingly become minutes-long tasks.
Together with Harmonic Security
Your employees are connecting AI to everything. Now what?
ChatGPT and Claude aren't just answering questions. Employees are connecting them directly to Notion, Linear, Jira, and the rest of your stack — with no security visibility into what data moves or what actions they take.
Harmonic Security gives your team the visibility to control it.
🔨AI Tools for the Shift
🌐 WebBrain – Adds an AI agent directly to your browser sidebar so you can research, understand pages, and complete everyday web tasks without constantly switching tabs.
🧠 Oasis – Creates a shared workspace where humans and AI agents can collaborate on real tasks, files, and projects together.
🔐 Execlave – Gives AI agents a secure gateway for connecting to external systems and taking real-world actions with tighter controls.
💻 Chiplab – Lets AI agents build, simulate, test, and debug embedded firmware on virtual chips without needing physical hardware.
🎥 Scrimba Explain – Turns questions into instantly generated video explanations, giving learners a more visual way to understand difficult concepts.
🚀Quick Shifts
👋 Microsoft is removing Mico, its expressive Clippy-like avatar, from Copilot voice mode less than a year after launch. Mico will move to Learn Live as Microsoft shifts Copilot toward a more unified assistant experience.
⚡Gemini Spark now runs on Gemini 3.7 Flash, improving how its 24/7 agent handles Workspace tools and complex multi-step tasks. Google says the model also brings major gains to coding, web development, and knowledge work.
💰 Writer launched Palmyra X6 alongside an upgraded agent harness designed to complete complex tasks with fewer tokens. Its research found harness improvements cut costs by around 40% on average, suggesting smarter orchestration can matter more than constantly switching models.
🤝 IBM is partnering with OpenAI to bring GPT-5.6, Codex, and ChatGPT Work to major enterprise clients. IBM will train tens of thousands of consultants on OpenAI technology, targeting large-scale deployments across finance, government, telecom, and retail.
🎵 Suno Studio 2.0 adds MIDI, automation, built-in effects, and a context-aware AI chatbot. Musicians can now play an idea, then ask AI to extend melodies, generate parts, or mix tracks.
🤳What’s trending on socials

🎬 A Free AI Video Studio That Runs Locally - A free open-source AI video studio with nearly 8K GitHub stars runs entirely on your PC with just 6GB VRAM, supporting Wan 2.2, LTX-2, Hunyuan Video, and Flux.
🧠 Google DeepMind Says AI Is Learning Without Humans - Google DeepMind’s VP of Research says models can now create their own challenges, judge their own answers, and improve themselves, as human training data runs out. The bigger prediction is that code generation becomes effectively free, with humans eventually no longer reading most AI-written code.
📌 Pinterest Users Are Getting Fed Up With AI Slop - Pinterest was supposed to be the quiet corner of the internet for real homes, haircuts, recipes, and art references. Now users are complaining that AI-generated imagery is overwhelming discovery, making something as simple as finding authentic human inspiration increasingly frustrating.
💸 A Top Economist Says the AI Boom’s Math Doesn’t Add Up - Apollo Chief Economist Torsten Slok warns that AI’s most profitable infrastructure layers are being paid by model companies still burning capital. His concern is simple: today’s upstream profits depend heavily on investor funding continuing until real customer demand and ROI can support the spending.
🗣️ Sarvam Opens Its Voice Agents to Everyone - Sarvam AI has opened its Voice Agents platform to everyone after powering 350M+ enterprise conversations. Businesses can now build human-like agents that remember context, personalize conversations, orchestrate workflows, and improve toward specific business outcomes.
That’s all for today’s edition. See you tomorrow as we track down and get you all that matters in the daily AI Shift!
If you loved this edition, let us know how much:
How useful was today's edition? |
Forward it to your pal to give them a daily dose of the shift so they can 👇



Reply