Correcting a summary is worth the thirty seconds only if the correction outlives the summary. An edit repairs one account of one conversation. A correction that is kept becomes a rule, and changes every account written after it. Almost everything about whether this is worth doing turns on which of the two you are making.
- The gap between a general summary and the one you wanted is stable rather than random. Too long, dates buried, the ask at the bottom, the same way every Monday. A repeatable fault is a fault a rule can reach.
- There are two ways a system can learn from you and they cost very differently. Training on preferences takes tens of thousands of judgments. Keeping what you said takes one sentence.
- A correction can make things worse. In the study that measured this on email, feeding users’ own feedback back into the classifier lowered its accuracy by about five per cent, for reasons that had little to do with the users being wrong.
- So what to check is not whether a system says it has learned. It is whether you can read back what it thinks you asked for, and whether the next message from the same sender arrives in the new shape.
What it means for a summary to learn
Three quite different things get called feedback, and only one of them carries a rule.
The first is a signal. A thumbs up, a thumbs down, a star. It records that something was good or bad without recording anything about what, and the space of things you might have meant by a downward thumb is large: length, order, tone, a missing date, a name in the wrong place. A signal is the cheapest feedback to give and the least useful to receive.
The second is an example. You rewrite the summary the way you wanted it and the rewrite is kept as evidence. That carries much more, and carries it implicitly: somewhere in the difference between what you were given and what you wrote is the thing you objected to, along with a dozen incidental choices you meant nothing by. You shortened it, and you also happened to use the client’s initials. One of those was the point.
The third is an instruction. One sentence saying what you want instead. “Lead with the dates.” “Never drop the amount.” “For this supplier, tell me only what has changed since last week.” It is the only one of the three that arrives already in the shape of a rule, because you wrote it as one.
Underneath the three sits the question that decides whether any of this is worth your time. Ask it of every correction you are about to make: would I make this same correction again next week, on a different conversation, from a different person? If the answer is yes, you are describing a rule and it should be stored as one. If the answer is no, you are repairing one summary, which is a perfectly good thing to do and simply has no future in it.
Most software offers you the first kind of feedback and describes it as the third. That is the gap this guide is about.
Why the same summary is wrong every time
A summary is one reading of a conversation rather than the correct account of it, and the measured disagreement between two people about what matters in the same document is wide enough that a single right answer is not on offer. A plain summary of every email works through that properly and it is the premise here rather than the subject.
What follows from it is the useful part. If there is no correct summary, then a summary written to a general standard is not randomly wrong for you. It is wrong in a consistent direction, because the gap between a general reader and your desk is a stable thing. You get the same disappointment every week, and it usually falls on one of five axes.
Length. The account is faithful and three times longer than the decision needs. Or it is so short that you open the thread to find out what it meant, which costs you the whole saving.
What leads. The facts are all present and arranged as a story, so the deadline that governs your week is in the fourth sentence. Order is not a cosmetic preference. It decides what you learn in the two seconds before you move on.
What has to survive. Every business has two or three things that must never be dropped: the amount, the statutory date, the file reference, the name of the person actually asking. A summary that omits one of them is not slightly incomplete, it is unusable, and it reads as complete because the other four facts are correct.
Vocabulary. Your trade has words for things and a general summary uses the sender’s words instead. “The registration” means something exact in your office and nothing in the abstract.
The sender. The same shape does not suit every correspondent. A weekly report from a supplier wants only what changed. A first message from a prospective client wants the whole of what they asked for.
The arithmetic is what makes this worth doing. A fault that happens once is a nuisance you absorb without noticing. A fault that happens on every status report from one supplier, every Monday, for a year, is fifty occurrences of the same small tax, and each one is paid at the moment you are least willing to stop and do something about it.
Two ways a system learns from you
The research on teaching a summarizer what people want is mostly about the first of these, and it is worth knowing what it bought and what it cost.
In Learning to summarize from human feedback, presented at the 2020 Neural Information Processing Systems conference by Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei and Paul Christiano, the authors set out to improve summaries by optimizing directly for what people preferred rather than for a similarity score. Their method worked, and the result is genuinely striking: a model with 1.3 billion parameters trained on human feedback beat a supervised model ten times its size, preferred over the dataset’s reference summaries 61 per cent of the time against 43 per cent. Both of their feedback-trained models were judged better than the human-written summaries they had originally been taught to imitate.
Now the price. The dataset they released to get there holds 64,832 summary comparisons, collected from paid labellers working to a written standard. That is what it takes to move a general model toward general human taste.
Two things about that do not survive the trip to your desk. The first is volume. Nobody running a seven-person practice is going to produce sixty-four thousand judgments about their own mail, and if the threshold for being understood is a number like that, you will never reach it. The second is legibility. What a model learns from preference data lives in its weights. There is no list to read, nothing to correct, and no way to find out why last week’s summaries changed shape.
The second way a system can learn is older and much less impressive to describe. It keeps what you said. One sentence, stored in plain words, applied to the next thing that fits it. There is no threshold to cross, because a single instance is enough, and there is nothing hidden, because the rule is a sentence in your language rather than a number in somebody’s model.
That is the mechanism that fits one person and one mailbox. Its weakness is exactly its strength inverted: it only knows what you have actually said.
Which is why a third thing usually runs alongside it, learning from what you do rather than what you say. The best public account of what that ought to look like is Guidelines for Human-AI Interaction, presented at the 2019 conference on human factors in computing systems by Saleema Amershi, Dan Weld and eleven colleagues, who codified more than 150 scattered design recommendations into eighteen guidelines and then had 49 design practitioners test them against 20 popular products that use AI. Their thirteenth guideline is the one in question: “Learn from user behavior. Personalize the user’s experience by learning from their actions over time.”
Watching what you open and archive is cheap, because you are doing those things anyway. The trouble is that behaviour is ambiguous in a way a sentence is not. You archived that thread unread because it was junk, or because you had already settled it on the phone, or because you were on a train. The archive looks identical in all three cases. Their fifteenth guideline exists for that reason: “Encourage granular feedback. Enable the user to provide feedback indicating their preferences during regular interaction with the AI system.” Behaviour is the cheap signal and the sentence is the clear one. A system that only has the first is guessing about you from evidence that does not distinguish between three different reasons.
What to correct and how to say it
Five kinds of correction are worth turning into a standing rule, and they map onto the five axes above.
- The shape. How long, how dense, whether it comes as a paragraph or as lines you can run your eye down.
- What leads. The one thing that should be in the first six words, every time, whatever else follows.
- What must always survive. The amounts, the dates, the reference, the name of whoever is doing the asking.
- The words. What your business calls the thing, in place of what the sender called it.
- How one sender is handled. The supplier who sends the same report weekly, the client whose mail is always about one matter, the colleague who buries the ask in the last line.
Three kinds are not worth turning into a rule. A wrong fact is a repair, not a preference, and the fix for it is to open the thread and read what was actually written. Anything you would not ask for a second time is a repair too. And a dissatisfaction you cannot put into words has nowhere to go: if you cannot say what you wanted instead, there is no rule in it yet, only an objection.
Then there is how to phrase it, which matters more than people expect.
- Say what you want instead, not what was wrong. “Too long” is a complaint and can be satisfied by cutting anything. “Give me the deadlines and the amounts first, then two lines of context” is a rule, and it can only be satisfied one way.
- Name the scope out loud. This is the correction people get wrong most often. A rule can apply to one sender, to one kind of mail, or to everything you read, and those are three very different instructions. Saying “shorter” while looking at a newsletter, and meaning newsletters, is how client threads end up truncated.
- One rule at a time. Two changes in one sentence give you nothing to test, because when the next summary is still not right you cannot tell which half failed.
- Correct on a representative example. The thread in front of you is the evidence your rule will be built from, so make the rule on an ordinary Tuesday rather than on the strangest email of the month.
- Say it while you are looking at the failure. Not later, not from a settings page.
That last one is not a convenience. It is the finding of one of the oldest results in the field. In Paradox of the active user, a 1987 chapter in Interfacing Thought published by MIT Press, John Carroll and Mary Beth Rosson described what they called the production bias: “Their paramount goal is throughput. This is a desirable state of affairs in that it gives users a focus for their activity with a system”, they wrote, and on the other hand “it reduces their motivation to spend any time just learning about the system, so that when situations appear that could be more effectively handled by new procedures, they are likely to stick with the procedures they already know, regardless of their efficacy.”
They were writing about word processors, and the sentence describes your Tuesday exactly. A correction that requires you to stop working, find a preferences screen and understand its vocabulary is a correction that will not be made, however handsomely it would repay the two minutes. Carroll and Rosson are explicit that this is not a defect in people to be trained out of them. It is a property of how anybody learns anything while trying to get work done, which means it is the design’s problem rather than yours.
Where learning from your edits goes wrong
There is one study that took ordinary people’s corrections to an email system, fed them back into the machinery, and measured what happened. It is worth reading closely, because the answer was not the flattering one.
Interacting meaningfully with machine learning systems, published in the International Journal of Human-Computer Studies in 2009 by Simone Stumpf, Vidya Rajaram, Lida Li, Weng-Keen Wong, Margaret Burnett, Thomas Dietterich, Erin Sullivan and Jonathan Herlocker, ran three experiments on a system that sorted email into folders and explained its reasoning as it went. The first was a think-aloud study: participants saw why the system had filed a message where it did and were free to say anything they liked about it, with no restriction on the form of their feedback. They gave rich, specific, sensible advice, which was the encouraging result.
The second experiment put that advice into the algorithm as constraints, and accuracy went down. The constraint-based approach came out below plain online training by five per cent when it used the participants’ keyword feedback, and by around three per cent on the other paradigm. The authors went looking for why, and their three reasons are ordinary enough to recognise in your own working week.
Participants proposed the most obvious word for the folder, which the system had already learned, so the feedback was redundant and bought nothing. The handful of words they offered was too small a fraction of what the system was weighing to change any decision. And some of the advice, taken literally, “degraded the classifier by over-constraining the parameter space”, which is a technical way of saying a true statement can still be a harmful rule if it is applied more rigidly than you meant it.
Their own conclusion is the honest one to carry away: “There are situations in which performance decreases by taking user feedback into account, possibly due to introduction of human errors into the learned information.”
The third experiment is the twist, though. Taking the very same keyword feedback through a different mechanism produced a highly significant improvement instead, at p equal to 0.0057. Same people, same corrections, opposite outcomes. Whether your correction helps is largely a property of how the system takes it up, and not of how well you phrased it.
Five failures show up in ordinary use, and all five are visible if you know to look.
Over-reach. One correction applied wider than you meant it. This is the commonest by a distance, and it is quiet: what breaks is not the thread you corrected but a different one next week.
The rule that was already true. You ask for something the system was doing anyway, nothing visibly changes, and you conclude that corrections do not work here. Stumpf’s participants did exactly this.
The stale rule. February’s instruction is wrong in June. Rules do not expire on their own, and a mailbox that has been learning for a year with nothing ever deleted is carrying preferences nobody would write today.
The silent change. Summaries start arriving in a different shape and you cannot tell whether that was your doing, somebody else’s update, or drift. Amershi’s group has three guidelines about precisely this, and they read like a specification: “Update and adapt cautiously. Limit disruptive changes when updating and adapting the AI system’s behaviors.” “Convey the consequences of user actions.” “Notify users about changes.”
The rule you would not have written down. Something learned from an unrepresentative moment, or from behaviour that meant nothing. Nobody would type “summarize everything the way I wanted it at 6pm on the worst Tuesday of the quarter”, and a system inferring from actions can arrive there without being asked.
Every one of those becomes manageable under the same condition, which is why it is the requirement to insist on. The preferences a system holds about you have to exist as a list you can read, in your own words, with each line editable and removable. Not for tidiness. It is the only way to tell an over-reaching rule apart from a system that is simply wrong, and those two problems have opposite fixes.
How to tell whether it learned
The test is cheap and it takes about two weeks of normal mail.
- Write down what you asked for, and its scope, in the words you used. Not a summary of your intent. The sentence, so you can hold the result against it later.
- Wait for the next message inside that scope. Same sender, same kind of mail. That is the only fair test. A different sender’s thread proves nothing either way, and judging a per-sender rule on a stranger’s email is how people talk themselves out of a feature that was working.
- Check the edge as well as the middle. Look at one conversation that sits outside the scope you named and confirm it did not change. A rule that improved the thing you aimed at and quietly altered everything else has not done what you asked.
- Count corrections, not summaries. The measurement that matters is not whether any one summary is right. It is whether you are making the same correction a second time. One repeat is a rule that did not stick or a scope that missed; three repeats is your answer about the system.
- Read the list once a month and delete what has stopped being true. This takes a minute and it is the only maintenance any of this needs.
Two boundaries on what the test can tell you. A summary will not converge on the one you would have written yourself, because that target does not exist for anybody. The honest goal is a summary wrong in ways you can live with rather than wrong the same way every week. And what you are measuring is a rate rather than an instance: some conversations will always confound a general rule, and the fifth one that does is not evidence that nothing was learned.
What a correction does not change
Four things sit next to this and none of them moves when you correct a summary.
Not the ranking. Saying that an account should be shorter says nothing about whether the conversation was worth reaching you first. Which threads surface, and on what evidence, is a separate measurement worked through in how email triage works, and the corrections that feed it are a different kind of statement about a different question. Whether that side of it can be trusted has its own test, run the same way and on different evidence.
Not the date or the task. A rule can insist that deadlines lead every summary. Turning what a conversation asks of you into something dated, that survives the thread being archived, is a different object with a different life, and email task management is where it lives.
Not the briefing. A preference about how one conversation is described is not a preference about how your morning is presented. They are separate reads answering separate questions, and a rule that suits one will often be wrong for the other.
Not what happened. A rule changes the account, never the conversation underneath it. The original messages are still what was actually said, they stay a click away, and no amount of tuning makes reading them unnecessary for anything you will be held to.
There is a fifth thing worth naming because people spend corrections on it. A summary can be badly shaped, which a rule fixes, or it can be strained by its source, which a rule does not. A conversation that has been running for six weeks and has reversed itself twice is hard to summarize for structural reasons, and those are set out in how to catch up on a long email chain. If the summaries you keep correcting are all of forty-message chains, the problem you have is not a preference.
How Point takes a correction
Point is a full email client running on the Gmail or Microsoft 365 account you already have, and every conversation in the feed arrives carrying a plain summary written by Point before you sit down.
When one of those is faithful and still not what you wanted, you say so where you are. Point puts the summary you actually saw beside the space where you write what you would change, so the correction is made against the thing that was wrong rather than from memory of it. One sentence in your own words is the whole of the input.
What happens next is the part that decides whether it was worth typing. The sentence is kept as a standing rule rather than spent on the summary in front of you. The next message of that kind, from that sender, is summarized fresh with the rule already in place. Tell Point that summaries for you mean action items with dates, and the following week’s status report comes back with three action items and three dates at the top instead of three deadlines buried in a story.
Every preference Point holds about you lives in one place, written as plain sentences rather than as settings, with the newest at the top and marked as added by you. You can write one directly without waiting for something to go wrong, change the wording of any of them, or remove one that has stopped being true. Nothing Point has learned about how you like things is held in a form you cannot read back.
The same rules carry into what Point writes for you. Ask for a reply and the sign-off has your phone number on it because you said once that client mail should, on a thread where you never mentioned it. A friend gets the register you asked for rather than the formal one Point began with.
How much Point does unprompted is a level held separately for each type of action, running from suggest-only through review to fully handled, and every type starts on review, so work arrives prepared and waits. Everything Point has done sits in one running log in plain language, and the entry that recorded an action is where you go to turn it back. The benefits page lists the rest of what is finished before you open the inbox, and what an AI email client is is the argument for the reading and the remembering living in the client rather than beside it.
Common questions
Does correcting an AI summary actually change anything?
It depends entirely on whether the correction is stored as a rule or spent on one summary, and those look identical at the moment you make them. The way to find out is to name what you asked for, then check the next message from the same sender that falls inside the scope you named. If the shape has changed there and nowhere else, the correction was kept. If you find yourself typing it again a fortnight later, it was not.
How many times do I have to say the same thing?
Once, if what you said is being kept as a standing rule. The reason to insist on that number rather than treating repetition as normal is the arithmetic in this guide: a fault that recurs weekly on one sender’s mail is fifty occurrences a year, and a system that needs to be told fifty times has not learned anything. Making the same correction twice is the finding, not the process.
Will one correction change every summary I get?
That is the risk, and it is the one worth being deliberate about. A correction can properly apply to one sender, to one kind of mail, or to everything, and stating which you meant is the difference between a fix and a new problem. The most common way this goes wrong is a rule made while looking at an unimportant email and applied to the ones that matter, so check a conversation outside the scope after you make a rule, not just the one inside it.
Can it work out what I want without me telling it?
Partly, and less than you would hope, because what you do is ambiguous in a way a sentence is not. Opening, archiving and setting aside are all cheap signals and every one of them has three plausible readings. That is why the design guidance for systems like this asks for both: learning from behaviour over time, and a way to state a preference during ordinary use. A summary tuned only from what you clicked is being written by something inferring your taste from evidence that cannot distinguish between junk and already handled.
Is my correction used to train the AI?
That is a separate question from whether it is remembered, and worth keeping separate, because a preference held for your account and material fed into model training are very different arrangements with very different consequences for client confidentiality. What Point does with your mail on that front is set out in keeping your data out of AI training, which is the page to read before you decide.
What if the summary is wrong about a fact rather than badly shaped?
Then you are looking at a repair rather than a preference, and the fix is to open the thread and read what was written. A rule cannot make a mistaken account correct, because there is no general instruction that means “be right about this client’s deadline”. Repairs are worth making and they have no future in them, which is the test this guide runs on: would you make this same correction again next week, on somebody else’s conversation.
What if I change my mind about a rule?
You should be able to remove it in the same number of seconds it took to make. This is not a small point of convenience. Rules do not expire by themselves and a preference set in February can be quietly wrong by June, so a list you can read and prune is what stops a year of corrections turning into a shape nobody chose. If you cannot see what a system believes about you, you cannot tell a rule that over-reached from a system that is simply mistaken.
The short version
An edit and a correction feel the same to make and behave completely differently afterwards. One repairs the summary in front of you. The other becomes a standing rule and changes everything of that kind that arrives later, which is the only version that is worth interrupting your work for. The test that separates them is whether you would make the same correction again next week on somebody else’s conversation.
What is worth stating as a rule is the shape, what leads, what must never be dropped, the words your trade uses, and how one particular sender should be handled. Say what you want instead of what was wrong, name the scope out loud, and make one rule at a time. Then hold the system to two things: that you can read back what it believes about you in plain sentences, and that the same correction never has to be made twice. Corrections that vanish are worse than no corrections at all, because you paid for them.
What you end up teaching depends on what you read all day. An accounting practice mostly wants the statutory date and the figure at the front of everything; a consultancy wants to know what moved in the scope since Friday. If your mail looks like neither, the same morning is set out for other desks.