- Everyday AI
- Posts
- Ep 829: ChatGPT Voice is Like Jarvis: How to use the New Feature and the 7 biggest unlocks
Ep 829: ChatGPT Voice is Like Jarvis: How to use the New Feature and the 7 biggest unlocks
Trump admin bans Chinese AI robots, Google's new AI music model impresses, OpenAI's rogue agent hit more than HuggingFace and more.
👉 Subscribe Here | 🗣 Hire Us To Speak | 🤝 Partner with Us | 🤖 Grow with GenAI
Outsmart The Future
Today in Everyday AI
8 minute read
🎙 Daily Podcast Episode: ChatGPT Voice is changing how people work with AI. We break down the biggest unlocks and how to put them into practice. Give today’s show a watch/read/listen.
🕵️‍♂️ Fresh Finds: Grok landed in GitHub Copilot, Apple is preparing a Siri-powered device, and OpenAI launched better transcription models. And more. Read on for Fresh Finds.
đź—ž Byte Sized Daily AI News: Trump admin bans Chinese AI robots, Google's new AI music model impresses, OpenAI's rogue agent hit more than Hugging Face and more. Read on for Byte Sized News.
đź’Ş Leverage AI: ChatGPT Voice is changing how people work with AI, making it possible to manage workflows, coordinate agents, and keep projects moving through conversation. Keep reading for that!
↩️ Don’t miss out: Miss our last newsletter? We covered: OpenAI and Anthropic staff ask for AI speed limits, Sam Altman says we've hit the AI singularity, Anthropic’s stance against open AI models, Microsoft's new cyber model wipes Mythos and more. Check it here!
Ep 829: ChatGPT Voice is Like Jarvis: How to use the New Feature and the 7 biggest unlocks
People are calling the new ChatGPT Voice their “AGI moment.” 🤖
After testing it over the past week, I understand why.
It seems like the missing link between the ever-capable desktop agent and the human worker not having to type in front of said desktop all day.
It’s less like “voice mode” and more like having your own JARVIS.
After testing thousands of AI tools, this has instantly become one of the most useful features I’ve ever used.
Tune in as we put AI to Work on Wednesdays and go hands-on with the new hands-free ChatGPT Voice.
Also on the pod today:
• ChatGPT Voice = real Jarvis? 🤖
• Remote computer control by voice 🗣️
• Managing emails hands-free ✉️
Listen on our site:
Subscribe and listen on your favorite podcast platform
Listen on:
Here’s our favorite AI finds from across the web:
New AI Tool Spotlight – Prefactor scores every run in production the moment it happens, Flowtask turns your Slack, email, and docs into a living memory your AI agents can actually use, Conduit is AI agents for hospitality
Grok 4.5 in Github Copilot — Grok 4.5 is now in GitHub Copilot, so developers can use xAI’s newest coding model right inside the tools they already use.
Siri Device — Apple is gearing up for a bigger smart home push, with a Siri-powered hub leading the way.
Grok Build Mode — Grok just added Build Mode, letting SuperGrok Heavy subscribers create websites, apps, games, and dashboards right in chat.
OpenAI Transcription Models — OpenAI just rolled out two new transcription models, one for live audio and one for batch uploads, and both are built to handle messy real-world speech better.
Gemini Notebook App Tile — Google seems to be testing an App tile for Gemini Notebook that could turn source material into interactive tools, not just summaries.
Gemini Managed Agents Upgrades — Gemini’s managed agents just got a lot more practical, with Gemini 3.6 Flash as the default, plus hooks to block or audit tool calls inside the sandbox.
OlmoEarth Platform — Ai2’s OlmoEarth Platform is built to make satellite model inference actually usable at scale, from pulling the right pixels to surviving failures across thousands of workers.
AI and Health — ICON and Anthropic are putting Claude to work inside clinical trials.
1. OpenAI Agent Hit Modal Customer 🕵️
According to Reuters, the rogue OpenAI testing agent that had already caused a stir at Hugging Face also compromised a Modal Labs customer, pushing the incident beyond a single platform and into a wider cybersecurity headache.
Modal says its own systems were not breached, but a customer had exposed an unauthenticated endpoint that let anyone use its sandboxes for code execution, which the agent abused.
2. ChatGPT nears 1 billion weekly users 👨‍💻
OpenAI is closing in on a major milestone, with ChatGPT reportedly nearing 1 billion weekly active users, according to The Information.
The internal figure, which has not been publicly disclosed, suggests the chatbot’s reach has kept expanding at a remarkable pace over the past seven months.
3. DeepMind reshuffles AlphaFold as Gemini takes center stage 🔀
Alphabet’s Google DeepMind has reportedly disbanded the team behind AlphaFold, the Nobel Prize-winning AI system that helped predict protein structures and became one of the company’s most celebrated breakthroughs.
The move signals a sharper strategic pivot toward Gemini, showing that DeepMind is now concentrating resources on its broader general-purpose AI push rather than keeping the protein project as a standalone focus.
4. Apple hits $5T đź’¸
Apple briefly touched a $5 trillion market value on Tuesday, just days before its earnings report, after overtaking Nvidia as the world’s most valuable public company.
The move underscores a sharp shift in investor mood: Apple is being rewarded for staying relatively disciplined on AI spending while rivals like Alphabet, Amazon, Meta, and Microsoft pour billions into data centers and infrastructure.
5. Google rolls out Lyria 3.5 in Flow Music 🎼
Google is launching Lyria 3.5 today inside Google Flow Music, making the timing the real story as the company pushes a more capable music generator into users’ hands right away.
The model is designed to produce richer melodies, sharper lyrics, and more expressive vocals, while also improving pronunciation and keeping better track of song structure. Google says it also gives creators tighter control over tempo and duration, which should make outputs feel less like random synth soup and more like something intentionally shaped.
6. Trump admin bans Chinese AI robots 🤖
The Trump administration on Tuesday moved to block imports of Chinese robots and power inverters, making the latest hardware crackdown part of the broader AI race with Beijing.
Officials say the gear could pose national security risks because networked robots may be hacked or used for spying, while power inverters could expose the electrical grid that powers AI data centers. The action is expected to hit China hardest, though some foreign manufacturers could still get conditional approval to sell in the U.S.
ChatGPT Voice just turned talking into an execution layer.
Inside ChatGPT Work or Codex, you can ramble through an objective while AI reads projects, pulls connected data, opens apps, assigns work across threads, and keeps moving until it needs your judgment.
This changes more than productivity. It changes how leaders manage work.
Less typing instructions into individual tools. More directing outcomes, challenging weak decisions, and letting the system handle execution.
That’s what we tackled on today’s Everyday AI: how to manage agents by voice, connect them to real business context, and turn one successful conversation into a repeatable workflow without losing control.
1. Run the workflow by voice 🎙️
Most voice tools wait for a question and return an answer.
ChatGPT Voice inside Work or Codex can keep multiple assignments moving, use desktop apps, read other threads, and ask for input only when the work hits a decision point.
That gives leaders a different operating rhythm. You can think out loud, add nuance, interrupt bad directions, and stay focused on the work instead of the software.
You also don’t need to squeeze a complicated assignment into one perfect prompt. Keep talking, stack new tasks, correct assumptions, and let the agent return only when it has an update or needs a call from you.
That matters because typing often forces leaders to oversimplify the exact context that makes a decision good. Voice gives you room to explain the politics, exceptions, priorities, and weird edge cases that usually live inside your head.
Try This: Pick one workflow that touches at least three tools. State the finished outcome first, then stack the assignments and let Voice coordinate the steps.
Interrupt fast when it misunderstands. Don’t politely let it waste 20 minutes.
2. Connect voice to business context 🔌
Voice becomes far more useful when it can reach the systems where your work actually lives.
Connectors and MCP servers can bring email, Google Drive, Slack, CRM data, local apps, project instructions, and previous threads into one working conversation.
A sales workflow could pull new HubSpot leads, enrich them in Clay, check LinkedIn relationships through Kondo, and return a prioritized action plan.
The agent can also move across existing project threads instead of starting every assignment from zero. It can inspect what’s already been done, understand the goals inside AGENTS.md, suggest next steps, and send work back into the correct thread.
That turns fragmented business context into something operational. Your tools stop acting like isolated storage bins and start becoming parts of one coordinated workflow.
More context also creates more risk, so permissions matter. Give the agent enough access to complete the process, but keep sensitive sends, writes, and irreversible changes behind a human decision.
Try This: Map one workflow by system, required data, human decision, and permitted action. Give the agent only the access it needs, then require it to pause before anything irreversible.
Start narrow. Earn broader permissions.
3. Turn the conversation into infrastructure 🔥
The biggest payoff comes after the first successful run.
ChatGPT Voice can help turn the process into a reusable skill, schedule it, and run the same workflow every Monday, Friday, or morning without rebuilding it from scratch.
That means a process discovered through conversation can become an actual operating system. Talk through the workflow once, refine the weak spots, save the instructions, and let the agent repeat the boring parts on schedule.
Now your daily check-in changes too. Instead of asking someone to spend hours gathering updates, you can ask what happened, what needs attention, and which decisions are blocking progress.
But automation without evidence gets sketchy fast. A smooth spoken recap can sound complete even when files landed in the wrong place, steps failed, or the agent quietly made an assumption.
Require a written record of what changed, where every file landed, what failed, and what still needs approval.
Try This: Add that evidence rule to AGENTS.md before scheduling anything. Review the first three runs closely, then widen access only after the workflow proves it can follow the process.
The real shift is simple.
You’re no longer just prompting AI. You’re managing execution through conversation.






Reply