// 9TO5GOOGLE — MOBILE & WEB
Everyone’s talking to their phones again
Good morning and happy Thursday. Today I’m donating to Dolly Parton’s Imagination Library in her honor, reading about OpenAI’s incredibly powerful new AI inference chip, pretending to read Bill Gates’ 6,000-word essay about AI’s diverging futures, and actually watching this 2-hour deep dive of the upcoming Thief: The Dark Project Remastered instead.
As a reminder, Inbox is the tech newsletter that goes beyond the news cycle. We cover the most important 9to5Google stories from the last couple of days along with a few highlights from around the web. We publish every Tuesday and Thursday, and if you like what you read, please subscribe.
Inbox is written by Daniel Bader. Read previous issues here, and then catch up with Ben Schoon’s excellent Weekender column.
Is talking the new typing? Despite a persistent trend towards resurrecting physical keyboards on smartphones, improvements to voice-to-text AI models have set up a parallel future where we hardly touch our devices at all anymore.
Purpose-built dictation and transcription apps have also taken off. I’ve cycled between Wispr Flow, Yaps, and Typeless on mobile, and Aqua Voice and Monologue on my Mac, and while they all have their advantages and quirks, they ultimately do the same thing. And like so many SaaS products, they run on a freemium pricing model that provides a small amount of free usage and an $8 to $12/mo fee for premium features.
Last month, OpenAI unveiled GPT-Live, a model that allows close to real-time full-duplex communication and sits on top of a multimodal LLM it can delegate tasks to. It was the latest in a long line of products focused on overhauling the reliability of voice inputs on any device with a microphone, from laptops to desktops to phones.
Now, Google has (re)entered the fray with Gemini 3.5 Transcribe, a model that promises to be faster and more accurate than GPT-Live and, if you’re a Pixel 11 owner, is already integrated into the excellent Rambler tool that’s Sherlocking many of the above products.
It can do all the things GPT-Live can, just faster. And unlike OpenAI’s, Google’s digital canvas is enormous. It wants you to talk to your Google Docs and Sheets, Gemini Live and Spark on Mac and mobile, and potentially more disruptively, through its Gemini API, by which developers can integrate voice input into their own apps.
Voice input isn’t new or novel — I’ve been sending WhatsApp voice memos to friends and family for years — but this new paradigm combines accurate transcription, where a model intelligently removes verbal tics and reformats text into tables or charts, with dynamic action. Think dictating an email that includes instructions to insert a specific photo from your camera roll and embed a map with directions. The combination of speed and dynamism is certainly enticing, but I’m not yet a complete convert.
There’s something tactile and contemplative about typing and communing directly with a thought that can’t be recreated with voice. And a world where everyone on the subway is dictating their commands to an agent sounds awful. But there’s a growing contingent of people, particularly those with fine motor impairments, who will find massive benefit from intelligent voice models, and that’s a world I’m OK with.