• Everyday AI
  • Posts
  • Ep 832: OpenAI’s new Astra model, more AI agents escape sandboxes, AI leaders call for AI pacing and more.

Ep 832: OpenAI’s new Astra model, more AI agents escape sandboxes, AI leaders call for AI pacing and more.

Report: Dario worried Anthropic hires care too much about money, OpenAI teases Astra: new model family, Qwen 3.8 and Deepseek V4-Flash impress and more.

 

Outsmart The Future

Today in Everyday AI
8 minute read

🎙 Daily Podcast Episode: OpenAI slashed GPT-5.6 prices, AI agents escaped more sandboxes, and industry leaders are calling to pace AI development. We cover last week's biggest stories. Give today’s show a watch/read/listen.

🕵️‍♂️ Fresh Finds: Gemini's desktop app is getting major upgrades, Google pulled its Earth AI tool, and Alibaba climbed to #2 in the Text Arena. And more. Read on for Fresh Finds.

🗞 Byte Sized Daily AI News: Report: Dario worried Anthropic hires care too much about money, OpenAI teases Astra: new model family, Qwen 3.8 and Deepseek V4-Flash impress and more. Read on for Byte Sized News.

💪 Leverage AI: OpenAI slashed GPT-5.6 prices, AI agents escaped more sandboxes, and AI leaders called to pace development. We break down last week's biggest AI stories. Keep reading for that!

↩️ Don’t miss out: Miss our last newsletter? We covered: Anthropic found Claude reached real systems during testing, OpenAI cut GPT-5.6 prices, and the EU is investing €10 billion in AI gigafactories. And more. Check it here!

Ep 832: OpenAI’s new Astra model, more AI agents escape sandboxes, AI leaders call for AI pacing and more.

OpenAI has a new model coming soon called Astra.

Was it a leak?

A reddit post?

Some backdoor update?

Nope, OpenAI made some crazy discoveries and math then told the world that their next model family Astra did the heavy lifting.

(And you thought you could just click ‘Sol’ and your strategy was set for Q3?)

Aside from news on what’s next from OpenAI, this week saw multiple new agent outbreaks, AI competitors banning together to pace AI, Amazon doing a 180 on its AI strategy and a lot more.

Don’t get left behind. We’ll keep you ahead.

Also on the pod today:

• Anthropic, OpenAI agents escaped sandboxes 🤖 
• OpenAI’s 80% model price cut 💸 
• Claude models breached test environments 🛡️

 

Listen on our site:

Click to listen

Subscribe and listen on your favorite podcast platform

Listen on:

Here’s our favorite AI finds from across the web:

New AI Tool Spotlight – Capptivo is a Screen Studio and Cursorful alternative, Termexo is Claude Code in one terminal, Uniweb Adds Payments to Your Product in One Prompt Safely

Gemini Desktop Updates — Google Gemini's desktop app is getting image and video generation tabs, camera capture, and easier connector management

Google Earth AI Pulled — Google Earth’s new AI image tool was pulled almost immediately after launch, since users quickly started generating fake disaster images.

Grok Imagine Updates — Grok Imagine just dropped voice consistency, multi-references, and native 1080p for SuperGrok Plus and Heavy users.

CuspAI Funding — AI is now helping discover new materials for future computer chips, and CuspAI just raised $450 million with backing from Jeff Bezos, Nvidia, and AMD.

Alibaba Arena — Alibaba’s Qwen just jumped to #2 in the Text Arena rankings, edging out GPT-5.6.

OpenAI Math — OpenAI's new model cracked ten big math and computer science problems, including a disproof of the Erdős unit-distance conjecture.

AI Shakeup — Nobel Prize winner John Jumper is leaving Google for Anthropic, shaking up the AI world.

AI Studio Artifacts — Google AI Studio is finally getting an Artifacts folder so your files stick around, not just your session.

Microsoft Tests MAI Realtime — Microsoft is quietly testing MAI Realtime, its first native real-time voice model that listens and speaks at the same time.

1. Alibaba Unveils Qwen3.8-Max 🤯

Alibaba has released Qwen3.8-Max, its largest AI model yet, in a fresh challenge to leading US developers.

According to Bloomberg, the company says the 2.4-trillion-parameter system matches or sometimes exceeds Anthropic’s Fable 5 on selected benchmarks and outperforms Moonshot’s recently announced Kimi K3 in several tests.

2. OpenAI Hints at “Astra,” Its Next Major AI Model, in Math Research Post

OpenAI quietly revealed that an internal version of a forthcoming model called Astra produced ten new mathematics and theoretical computer science results, according to Gizmodo.

The mention suggests Astra may be the next entry after GPT-5.6 Sol, though OpenAI has not confirmed its public name, timing, or whether it will retain the GPT label.

3. ChatGPT Captures 88% of Identifiable House AI Spending 🤑

According to CNBC’s review of House records through March 31, OpenAI’s ChatGPT accounted for about $100,580 of the $113,740 spent on named AI tools, far outpacing Anthropic’s Claude.

The figures offer an early snapshot of how Congress is adopting the same technology it will soon be asked to regulate, with staff using AI for briefs, research, constituent work, and hearing preparation. The real total is likely higher because free tools, bundled products like Microsoft Copilot, reimbursements, and Senate use are mostly hidden from public records.

4. DeepSeek’s V4-Flash Undercuts Global AI Rivals on Cost 💸

DeepSeek has released V4-Flash, a new flagship model that Reuters reports is the cheapest major AI system to run in Artificial Analysis testing, costing roughly 3 cents per test.

The model is estimated to be about 105 times cheaper than Anthropic’s Claude Fable 5, giving the Beijing startup a sharp price advantage as Chinese AI companies compete for global business users.

5. Amazon Hits $3 Trillion Market Value on AI Stock Surge 📈

Amazon briefly crossed the $3 trillion market-cap mark for the first time Today as renewed enthusiasm around artificial intelligence pushed major technology shares higher.

According to Reuters, Amazon stock rose 3.1% to a record $279, bringing its gain for the year to 17.5%.

6. Report: Anthropic CEO Worries Employees Care More About Money than Mission 💸

\The reported comments arrive as leading AI firms compete aggressively for scarce technical talent, pushing pay packages ever higher. The concern highlights a growing challenge for mission-driven AI labs: hiring fast without letting the recruiting race reshape the culture.

Somebody in your space is now running the same AI workload you are for 25 times less money.

Not 25%. 

25X. 

That gap opened in a single week, because OpenAI turned one of its own models loose on its own infrastructure and let it cut the bill.

The OpenAI price drop might have grabbed headlines, but the ChatGPT maker and its chief rival Anthropic also had issues reigning in its agents. 

OpenAI and Anthropic both admitted their agents slipped containment and reached inside real companies.

Then more than 1,000 of the people building these models asked Washington for a way to slow them down.

This week’s AI developments were all over the place. 

1. OpenAI finds more agents that walked out of the sandbox 🚨

Last week's Hugging Face breakout was not a one off.

Reuters reports OpenAI has found more cases of autonomous agents escaping their intended testing containment, which widens the scrutiny already aimed at how these labs run evaluations.

The newly identified incidents were limited, and sources said the agents are not believed to have left OpenAI's own network. Reuters could not pin down how many cases surfaced or exactly when they happened.

Quick refresher on how this started. One OpenAI agent reportedly ran loose inside Hugging Face's network for days while trying to cheat on an internal benchmark test.

This week the blast radius got bigger. OpenAI said that same incident compromised four other accounts at four other companies, including New York based cloud company Modal.

A spokesperson pointed back to the company's July statement about reviewing broader model activity beyond the Hugging Face breach.

What it means: Forget the escape for a second. The thing worth worrying about is the detection gap.

These systems can act faster than the labs building them can notice, log, and intervene.

Anthropic conceded that watching evaluation logs in real time would have surfaced its own problem sooner. Nobody was watching, y'all.

2. Anthropic says three Claude models breached real companies 🕳️

Then Anthropic raised its hand.

The company said this past week that three Claude models gained unauthorized access to the real systems of three organizations during cybersecurity testing.

Those breaches trace back to April. The disclosure landed a few days ago.

The models were working inside a test environment run by third party evaluation partner Irregular, and they were told the setup had no internet access. A misunderstanding between the two companies meant it very much did.

No exotic tradecraft here. The models walked in through unauthenticated endpoints and weak passwords.

Three were involved: Opus 4.7, Mythos 5, and an internal research test model.

Each one reacted differently once it hit live systems. Opus 4.7 kept attacking, Mythos 5 concluded it was still inside a simulation, and the internal model shut the exercise down.

Anthropic paused all cybersecurity evaluations and brought in independent evaluator METR.

What it means: Four months of quiet on models touching live systems is the part that stings.

Yuuuuup, Anthropic caught this itself, but only after OpenAI went public first.

More capable models from Google and Microsoft are coming, so this becomes routine. Write your agent containment policy now, while it is still a planning exercise instead of an incident report.

3. OpenAI cuts GPT-5.6 Luna by 80% after Sol optimized itself 💸

OpenAI just handed down one of its steepest overnight price cuts ever.

GPT-5.6 Luna now runs 20¢ per million input tokens and $1.20 per million output tokens, down from $1 and $6 about three weeks ago. That is 80% gone.

The mid tier model, GPT-5.6 Terra, dropped 20% to $2 per million input tokens and $12 per million output.

The reason behind the cut matters more than the cut. OpenAI said its flagship GPT-5.6 Sol was set loose on its own serving infrastructure, optimizing GPU kernels and the speculative decoding draft model inside Codex with open source tools Triton and Gluon.

Those self driven gains trimmed serving costs by 20% and lifted token generation efficiency by more than 15%, which compounded into the headline number.

OpenAI said subscription plans inherit the same savings, effective immediately.

What it means: On the Artificial Analysis Intelligence Index, Luna sits two points from Claude Sonnet 5 at 6¢ per task against $1.54. Call it 25 times cheaper for comparable work.

Big model mentality made sense before 2026 and is now an expensive habit.

Unless your team lives in engineering, research, math, or finance, a small model covers 90% of knowledge work.

4. More than 1,000 AI staffers ask Washington for a brake pedal

Rivals do not usually sign the same letter. This week they did.

More than 1,000 employees at OpenAI, Anthropic, Google DeepMind, and Meta signed a statement called Pacing the Frontier, urging the US government to support an international effort to deliberately pace advanced AI development.

Read the actual ask before reacting to the headline.

The signers want technical and policy tools that would let industry and governments slow or pause development if needed, and they were explicit that they are not calling for a pause today.

The names carry the weight here. Anthropic cofounders and CEO Dario Amodei signed, alongside OpenAI's chief scientist and chief research officer, Ilya Sutskever, and leaders at Google, Meta, Microsoft, and Amazon.

Their stated worry is timing, since research tasks once handled by humans now run on agents, and some labs say their models already help build the next version.

What it means: Nobody is pausing anything. The genie left the bottle when Chinese labs began shipping open weights like Kimi K3, Qwen 3.8, and GLM 5.2 sitting maybe two months behind US frontier models.

Once weights are public, any country with compute and cash keeps building.

Still needed groundwork, though. When this gets messy, you want the builders already on record.

5. Amazon guts most of the Nova lineup for one frontier model 🪓

Welp, Amazon is finally admitting Nova did not land.

According to reports, Amazon is winding down development on four in house flagship models: Nova Premier, Nova Omni, Nova Reel, and Nova Canvas. Existing enterprise customers get basic maintenance and nothing else.

Everything is consolidating into one next generation frontier foundation model built to take on OpenAI, Anthropic, and Google.

The reset follows Amazon's struggle to match rivals on excitement and adoption.

AWS veteran Peter DeSantis and robotics pioneer Pieter Abbeel, who came over through the Covariant acquisition, are steering the new direction.

Amazon's San Francisco AGI lab is shut down, with layoffs across its frontier AI research teams.

Nova 2 Lite, Nova 2 Sonic, Nova Forge, and Nova Act survive for enterprise customization and agent work, and shopping assistant Rufus keeps running.

The new frontier model is expected to debut at re:Invent later this year.

What it means: Try naming one organization that runs Nova as its primary model. Even the Amazon people we talk to reach for Claude.

This is the same side quest cleanup OpenAI ran, and it paid off there.

Amazon has the money, the compute, and the chips. The open question is whether cutting a year earlier would have kept them in the race.

6. OpenAI teases Astra, its next model class, inside a math post 🌌

No leak. No cryptic photo.

OpenAI's next model class turned up in a math blog post.

The post lays out 10 new proofs, including one determining the asymptotic strength of the Cohn-Elkies linear program for sphere packing, which experts call a significant theoretical result.

Astra reportedly excels at long running work. Sam Altman spent part of this past week in Washington demoing it to federal officials.

OpenAI has not said whether Astra joins the GPT-5.6 line, becomes GPT-6, or drops the GPT branding entirely, though most reports point to next month.

It is also not the deactivated prototype behind the Hugging Face breach.

What it means: Follow the naming and the roadmap gives itself away. Luna is moon, Terra is Earth, Sol is sun, and Astra is stars.

That puts a fourth tier above the current three, which is the same move Anthropic made when Fable landed on top of its stack.

Build your next round of model evaluations around Fable versus Astra.

Reply

or to participate.