Skip to content

Talk to your inbox when your hands are full

On this page

Most mornings there’s a minute when both hands are full. You’re carrying files in from the car, or standing at a printer, or making coffee before a 9:00. Mail lands in that minute anyway. In Point the minute is usable. Hold the mic, say what you want in an ordinary sentence, let go, and listen.

  • The mic sits beside the box you’d otherwise type into. It’s push to talk. Point hears you while your finger is down, and there’s no wake word to say.
  • Same assistant, same box, so anything you can type you can say. What changes is the answer. It comes back spoken and short, with the full version waiting on a card.
  • Speaking a message runs about three times faster than typing one. Listening runs at about the speed of reading, and that gap is the single most useful thing to understand about using this well.
  • The mic is a way in. How much Point finishes on its own is a setting you hold, and saying something out loud leaves that setting where it was.

The mic beside the box you already type into

In the feed, next to the box where you type an ask, there’s a microphone. Hold it, talk, and let go when you’ve finished the sentence. That’s the whole thing.

Push to talk is worth a moment. Your finger is the switch. Point hears you while you hold, and stops the second you release. The holding is the listening. There’s no phrase to say first, and nothing sitting open across the room waiting for its own name. If your office is also where clients sit down, that’s what makes this a feature you can leave switched on.

What goes in is what the typed box already takes, so everything you’ve learned there works out loud:

A name or two, to look something up. A client, a company, a person. Back comes what Point holds under that name. That’s One box that searches mail, tasks, calendar and files, reached with your voice.

A description, when the exact words are the part you lost. Finding an email when you cannot remember the words sets out how to describe a half-remembered message so it comes back.

A question whose answer lives inside a file. The total on the invoice, the notice period in the lease. That’s asking a question and getting the answer out of the attachment.

A sentence with a verb in it. Archive these, draft that, move Thursday. Saying what you want done in one ordinary sentence covers what makes one of those land. Every word of it holds for a spoken one.

What comes back is short, and the rest waits on the card

Say read my feed and two or three sentences come back, most urgent first. The top item, named, with what it needs from you. Then a count of what else is waiting, and roughly what kind of thing it is. On screen at the same time, a card carries the detail: the thread, the sender, the buttons. Nothing is lost in the short answer. The card holds the rest of it.

That split is the design, and it’s worth knowing why, because it tells you what to ask for out loud.

Speaking in is genuinely fast. In December 2017, Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies published a study comparing speech against a phone keyboard for short messages. Sherry Ruan, Andrew Ng and James Landay of Stanford ran it with Jacob Wobbrock of the University of Washington and Kenny Liou. “We found that with speech recognition, the English input rate was 2.93 times faster (153 vs. 52 WPM)”. Speech made fewer errors along the way too, a 5.30 percent corrected error rate against 11.22 percent for the keyboard. It left very slightly more in the finished text, 1.30 percent against 0.79 percent. The authors are careful about the conditions. This was an iPhone in a laboratory, and they say plainly that “further study is required to quantify performance in non-laboratory settings for both methods.”

Coming back out, the gain flattens. Marc Brysbaert pooled 190 studies covering 18,573 participants, in the Journal of Memory and Language in December 2019. He put the average silent reading rate for adults in English at 238 words per minute for non-fiction. He noted that “the reading rates are in line with maximum listening speed.” Listening keeps pace with reading. It’s simply linear.

That’s the real constraint. Your eye drops onto the fourth row of a list, skips the two you recognize, and reads the sixth first. Your ear takes one thing at a time. A spoken list of nine things arrives as nine things in a fixed order, and by the fourth you’ve lost the first.

So the working rule is: ask out loud for one thing, or for the top of something. What’s the most urgent thing this morning. What did Marcus say about the lease. Is anything waiting on me before the 2:00. Each of those has an answer. List everything unread from this week has an inventory, and an inventory belongs on the card where you can skim it. The card is the half of the answer that does the skimming.

Talk over it the way you would a colleague

Cut in whenever you like. Point is three sentences into an outage, you already know what you want done, so you say it. The sentence you interrupt with is the one that gets carried out.

Follow-ups keep the subject too. Ask about the outage, hear the summary, then say tell me more about that or who else is on the thread. The outage stays understood. You’re in a conversation, and a spoken exchange usually runs shorter than the typed one.

This is the part that changes the habit for most people. Having to hear an assistant out is what makes people use it twice and drop it. One you can cut off mid-sentence is one you keep, because the worst case costs you two seconds.

Wording it for the ear

The three things that make a request work are the same spoken or typed, and they’re set out next door rather than here. What the ear adds is a handful of places where talking and typing genuinely differ.

Say names the way they sound, then look at the card. Speak them as you would to a person. A name Point hears wrong shows up wrong on the card in front of you. That’s the cheapest place in the world to catch it, and the fix is a tap.

The punctuation you’d have typed becomes words. A comma before don’t send it is silent. Say and hold it, don’t send it yet and the limit survives the trip. This is the one habit that takes a week.

One hold, one job, unless you’d say it in a breath. Find Priya’s last note about the audit and remind me Thursday to answer it is a single natural sentence and works as one. Four clauses with an unless in the middle is a sign the job wants the screen. You’d feel the same saying it to a person.

Say the limit, because out loud you will anyway. The only this month. The put it on Tuesday, she’s out. Typing drops those, and talking brings them straight back, because that’s how you’d brief somebody covering your desk. Spoken requests come out the most complete of any.

Type the strings that have to be exact. An invoice number, a case reference, a policy number. More on that below, and it belongs in the habit list: those go in with fingers.

The microphone does not move the dial

This is the question underneath all of the others, so it’s worth answering flatly. Speaking is an input method. It is not a higher level of trust, a shortcut past a confirmation, or a different mode with different rules.

How far Point takes a piece of work before it reaches you is held per kind of work. Each kind runs from suggest only, through review, to handled outright. Out of the box, every kind of work sits in the middle. Setting how much your inbox does on its own is where those live, and how to think about moving one.

So say reply to Daniel that I’m on it tonight. With drafting on review, the reply comes back written and waiting in the card, with the same buttons a typed request gives you. Move that setting up and the same sentence behaves differently. The setting decides it, however you said it.

Two things follow. The first time Point would take a kind of action you haven’t seen before, that’s its own decision and it comes to you on its own, which is approving anything new before it acts. And whatever was done is written down the same way, typed or spoken. That matters most right here. A typed request leaves a sentence on your screen, and a spoken one fades as you say it. The activity log behind every action is where the record lives, and undoing what Point did is where its reach ends.

Speaking carries the same weight as typing. It’s simply the one input that leaves no trace of itself, so if you read that log in your first week or two, voice is the reason it’s worth it.

The car is the case worth being careful about

Say hands-free and most people picture driving. It’s the obvious use, and the honest answer here is a caution. So here’s the reading.

The National Highway Traffic Safety Administration’s research note on distracted driving in 2024, published in April 2026, reports that “in 2024 there were 3,208 people killed and an estimated additional 315,167 people injured in motor vehicle traffic crashes involving distracted drivers.” Eight percent of fatal crashes that year were reported as distraction-affected.

What hands on the wheel buy you is a separate question, and the research on it is specific. In October 2015 the AAA Foundation for Traffic Safety published Measuring Cognitive Distraction in the Automobile III, by David Strayer and colleagues at the University of Utah. In it, 257 subjects spent a week with one of ten model-year 2015 cars and their built-in voice systems. Three findings are the ones to carry.

Hands-free still costs you attention. Their conclusion is direct: “the higher levels of workload should serve as a caution that these voice-based interactions can be cognitively demanding and ought not to be used indiscriminately while operating a motor vehicle.”

Practice leaves the cost in place. “Practice did not eliminate the interference from IVIS interactions. In fact, IVIS interactions that were difficult on the first day were still relatively difficult to perform after a week of practice.”

And the cost runs on after you stop talking. “Significant residual costs were observed for 27 seconds after the IVIS interaction had terminated.” At the 25 mile per hour limit used in the study, that’s 988 feet of road covered before attention has fully returned.

All of that is a finding about driving while you talk, whichever assistant is on the other end. That’s exactly the point of quoting it. The cost sits in your attention, so better software leaves it exactly where it is. Whether your state allows it is a separate question from whether it’s a good idea, and state law varies.

The practical version is simple. Park first. The minute before you start the engine and the minute after you switch it off are both minutes when your hands are full and the road wants nothing from you, and that’s where most of the value of a hands-free ask actually sits.

Four times to type it instead

When the words have to be exact. An invoice number, a docket number, a clause you mean to quote. Speech adds a transcription step to a job with no tolerance for one. Exact strings already belong to letters rather than meaning, and that goes double when the letters are spoken. Where exact words still win has the general version.

When you can be overheard. A spoken client name carries across a coffee shop, a shared office or a client’s waiting room. It travels further than a typed one, and the mic has no idea which room it’s in. This is the oldest confidentiality problem there is, and it deserves a rule you keep rather than a judgment you make each time. Client confidentiality and AI covers the part that’s about the software.

When recognition keeps failing you. Background noise is the obvious cause, and moving three feet from the espresso machine fixes most of it. The other cause is worth naming plainly. Recognition works better for some voices than others. In 2020 the Proceedings of the National Academy of Sciences published a study by Allison Koenecke and colleagues. They tested five systems from Amazon, Apple, Google, IBM and Microsoft against interviews with 42 white and 73 black speakers. They found “an average word error rate (WER) of 0.35 for black speakers compared with 0.19 for white speakers,” and traced it to the acoustic models themselves rather than to what was said. That’s a finding about the field, across every product in it. So if a voice interface misreads you constantly, you’re saying it fine. Put the things that matter through the keyboard and speak the rest.

When you’re composing rather than asking. Speaking one sentence is fast. Four paragraphs of a difficult client email is a different job. It runs past you while you talk, and you can’t see it while you do it. Ask out loud for the draft, then read it and edit it with your eyes. That’s what the draft is for.

Common questions

How do I use voice in Point?

Hold the microphone next to the ask box in the feed, say what you want in an ordinary sentence, and let go. It’s push to talk, so Point hears you while your finger is down and stops when you lift it. Your first one works right now, with no voice mode to enter and nothing to configure first.

Is Point listening all the time?

No, and that’s the reason it works this way. There’s no wake word, so nothing sits waiting to hear its name. Your finger starts it and your finger ends it. Between uses, the microphone is nothing you have to think about.

Can I do everything by voice that I can do by typing?

Same assistant, same box. A lookup, a described memory, a question about an attachment, a request with a verb in it: all of them work spoken. What changes is the shape of the answer rather than what you’re allowed to ask. Anything whose honest answer is a long list reads better than it hears, and it’s waiting on the card.

Will it read my email out loud?

Point gives you the short version out loud, most urgent first, with the detail on screen at the same time. Ask read my feed and the top item comes back named, with a sense of what else is waiting. A few sentences, rather than a recital. Then ask about any of it without naming the thread again, and cut in whenever you’ve heard enough.

If I ask out loud for a reply to be sent, does it send?

That depends on a setting rather than on the microphone. How far Point takes each kind of work before it reaches you is held per kind of work, and it starts on review. So out of the box, a spoken reply to Daniel comes back as a draft waiting in the card, exactly as a typed one would. Speaking leaves that setting exactly where you put it.

Can I use this while I am driving?

The research says be careful, and it says it about voice in cars generally. The AAA Foundation’s 2015 in-car study found voice interaction cognitively demanding. Practice left the effect in place. Attention was still impaired for 27 seconds after the interaction ended. Their own recommendation is that voice-based interactions “ought not to be used indiscriminately while operating a motor vehicle.” Before you pull out and after you park are the better minutes, and state law on the question is separate again.

What if it mishears a name or a number?

You’ll see it on the card first, because the spoken answer and the card arrive together. Correct it there, before anything happens with it. For anything that has to be character-perfect, an invoice number or a reference code, the keyboard saves you fixing a transcription afterward.

Does it work in a noisy office?

Reasonably, and better with a few feet between you and the noise. If it misreads you often enough to be annoying, type the things that matter and speak the rest. Recognition quality varies by voice as well as by room. That’s a known property of the technology, and you’re saying it right.

The short version

The mic sits next to the box you already type into, you hold it while you talk, and there’s no wake word between you and it. Everything you can ask in writing you can ask out loud, because it’s the same assistant. The return trip is what differs. Speaking in runs roughly three times faster than typing, while listening runs at reading speed and costs you the skim. That’s why the spoken answer is two or three sentences and the card carries the rest. Ask out loud for one thing. Say the limit the way you would to a person. Interrupt when you’ve heard enough, and type the strings that have to be exact. Then hold the two boundaries that sit outside the software: the microphone leaves what Point is permitted to finish exactly where you set it, and a car is the wrong place to find that out. The rest of what Point will do when you ask it is on the benefits page, and most of it answers to a sentence whether you type it or say it.

Ready for a calmer inbox?

Join the private beta

We're onboarding a few teams at a time. Leave your email, confirm it once, and we'll send an invitation the moment a place opens.

By joining you agree to our privacy policy.

Private beta

What you're joining

It runs on the mail you have

Point sits on top of Gmail or Outlook. Your address, your history and your contacts stay exactly as they are, so there is nothing to migrate.

You set how much Point does

Out of the box everything waits for your review, replies included. You hand over only what you trust, one kind of work at a time.

Join the private beta

We're onboarding a few teams at a time. Leave your email, confirm it once, and we'll send an invitation the moment a place opens.

By joining you agree to our privacy policy.