Skip to content
← Guides

AI email triage, and whether you can trust it

Part of Email triage

On this page

AI email triage means letting a model read your unread mail and rank it before you do, so what needs you is already at the top. The method is old; reading all of it is new. The question is not how it works but whether the judgment is any good, and what happens the first time it is wrong.

  • A model is reliable on the judgment you would make in a second, and unreliable on the one that needs context it never saw.
  • Gmail’s Priority Inbox reached roughly 80% accuracy, and most of that came from learning one person at a time rather than from a cleverer model.
  • People abandon an imperfect algorithm faster than an equally imperfect person. Being able to correct it is what makes it usable.
  • Test it against your own inbox for two weeks. The check is cheap and the answer is specific to you.

What is AI email triage?

Triage is the pass that sorts arriving mail by what it will cost you to ignore it, and how email triage works covers the method in full. AI email triage is that pass run by a model rather than by you, over everything that arrived, before you sit down.

Search the term and you will mostly find developer tutorials: wire this webhook to that classifier, add a queue, ship it. Useful if you are building one, beside the point if you run a firm and want to know whether to let it near your client mail. That is the question this one answers, and who Point is for takes it trade by trade.

What can a model actually judge?

There is a good public answer, because Google published theirs. The Learning Behind Gmail Priority Inbox, a 2010 research note by Douglas Aberdeen, Ondrej Pacovsky and Andrew Slater, describes a system that “ranks mail by the probability that the user will perform an action on that mail”, with accuracy “approximately 80 ± 5% on a control group”.

Two things there are worth sitting with.

The first is the definition. They did not predict importance, they predicted whether you would act. Mostly a good proxy, since the mail you act on is usually the mail that matters. But where behavior and interest come apart, it follows the behavior.

The second is where the accuracy came from. Their error rate fell from 45% using one model for everybody, to 38% with a model per user, to 31% once each user’s threshold was personal too. As the paper puts it, “importance ranking is harder as users disagree on what is important, requiring a high degree of personalization”. Nobody can build you a good ranking out of general rules about email. It has to learn your firm.

The payoff was measured on Google’s own staff, and it is real rather than dramatic. Across roughly 2,000 employees running it, they “spent 6% less time reading mail overall, and 13% less time reading unimportant mail”, and were “more confident to bulk archive or delete email”. The confidence is the bigger win. Time saved is pleasant; not having to check is what changes a day.

What can it not know?

A model reads the mail. It does not know what you agreed on the phone yesterday, that this client has been quietly unhappy since March, that an invoice for this amount is routine for one client and alarming for another, or that the partner who usually handles this is away. It cannot see the commitments you carry that never touched your inbox.

So the boundary is this. A model is good at the judgment you would have made in a second from the sender, the subject and the first two lines, and it makes that judgment on all of it rather than on the top of the pile. It is unreliable exactly where you would have needed to stop and think, and those are the messages you were always going to open anyway.

There is also a weakness inside the inbox. Judgment built on how you have dealt with a sender has nothing to work from the first time that sender writes, so a first approach from someone new is both the message you can least afford to miss and the one the model knows least about.

Why the first mistake matters more than the accuracy rate

In Algorithm Aversion, a 2015 paper in Journal of Experimental Psychology: General, Berkeley Dietvorst, Joseph Simmons and Cade Massey found that people “more quickly lose confidence in algorithmic than human forecasters after seeing them make the same mistake”. Five studies, in the Wharton School’s lab and on Mechanical Turk, forecasting rather than sorting mail. Participants dropped the algorithm even after watching it beat the human. One visible error was enough.

Worth knowing about yourself before you start, because your assistant will misfile something in the first week and your instinct will be to conclude the whole thing is unreliable. A person doing the same job would misfile more.

The same researchers found the way out. In a 2018 follow-up in Management Science, where the choice was between an algorithm’s forecasts and the participant’s own, people were “considerably more likely to choose to use an imperfect algorithm when they could modify its forecasts”, and that preference “held even when participants were severely restricted in the modifications they could make”. It was not more control they wanted. Some was enough.

Which gives you the test to apply to anything you are considering. Not how accurate is it, but what happens when it is wrong, and can I fix it in one move.

How much do you hand over, and how do you check it?

Not all of it at once, and not the same amount for every kind of work. Point puts a separate setting on each kind of task, running from suggest-only, through waiting for your review, to handling it outright. Review is what you get out of the box. At the top of a setting Point does that work without asking, which is why the settings move one at a time.

Sorting and filing is usually the first thing people hand over. Anything that goes to a client tends to stay on review far longer, for a reason worth stating plainly. Filing is reversible: one click on the record puts the thread back in your feed. A sent reply is not reversible by anyone, because it is already sitting in somebody else’s mailbox. Undo covers what Point did to your inbox and stops there.

Everything done for you is written down in order, in plain language, and corrections go in the same way rather than through a settings screen. Tell Point once that a summary buried the deadlines and that becomes a standing rule. Star a contact and everything they send scores higher from then on, which is your judgment entered directly rather than a guess the model has to make. And you can always audit it: one tap switches the feed from ranked to plain time order, everything in the sequence it arrived, nothing hidden. A ranking you cannot check is one you take on faith, and nobody should take an inbox on faith.

The trust question starts before the ranking, of course, particularly in a practice whose mail is full of other people’s financial affairs. Point holds suspicious mail at the door unread and unprocessed, in a quarantine only you can release, and individual sensitive messages can be locked end to end so only you and the recipient can read them. The security questions worth settling first go through the whole checklist, and the benefits page covers what you get in return.

How do you test it before you trust it?

Two weeks is enough, and you can run it without risking anything. Choosing between products is a different exercise, covered in choosing email triage software. This one asks a narrower thing: you have the tool, and you want to know whether to believe it.

  • Keep the ranking on and the acting off. Let it sort, leave everything else on review. You are testing judgment, not autonomy.
  • Read the top of the list, then read the rest anyway. For the first week, do both. You are hunting for a message it put low that you would have put high.
  • Count the misses, not the hits. A wrong ordering that cost you nothing is noise. One important message ranked low is the number that matters, and it should be zero after week one.
  • Correct rather than tolerate. Star the people who always matter, say when a summary is off. A system you correct is the one you end up trusting.
  • Read the log once a week. If the record is thin or hard to follow, that is your answer about the tool.

If the important mail is reliably at the top and the log holds no surprises, you have your evidence. If not, you have lost two weeks of reading your inbox twice.

Common questions

Is AI email triage accurate enough to rely on?

The best public figure is around 80%, from Google’s Priority Inbox, and personalization is what moved it. Not accurate enough to stop reading your inbox. Easily accurate enough to stop ranking it yourself. The measure that matters is whether anything important was ranked low, which you can check in two weeks.

Will it learn how my firm works?

That is where the accuracy comes from. Google’s result was that a model per user beat a general one by a wide margin, and personal thresholds beat that again. Expect roughly right in week one, noticeably better by week three, and expect to have corrected it a few times in between.

Can I still see my inbox normally?

At any moment. Point keeps a plain chronological view one tap away, and connecting Gmail puts the same three groupings on your actual mailbox as labels, so the judgment is not locked inside one app you have to remember to open.

What if I disagree with how it ranked something?

Say so, and it should hold. A correction you have to repeat is a bug, not a preference. Being able to modify an imperfect system is what makes people willing to use it at all.

The short version

A model can rank your mail about as well as you would at a glance, across all of it, before you arrive. It cannot know the parts of the job that email never touched, and it will get something wrong in the first week. What decides whether it is worth having is not the accuracy number but whether you can see what it did, disagree in one move, and have that stick. Run two weeks with the sorting on and the acting off, and count the important messages it put low. That number is your answer.

Join the private beta

We're onboarding a few teams at a time. Leave your email, confirm it once, and we'll send an invitation the moment a place opens.