Skip to content

When an email is written for the AI and not for you

On this page

Email has always carried messages written to fool the person reading them. Connect an AI assistant, and your mailbox carries a second kind, written to fool the software that reads first. The two problems have different answers. The second kind runs on arrival, with no click from you. In the worst documented case it ran before the message was opened. What holds it back is a filter you already pay for and a setting you already own.

  • The attack is dull to describe and close to invisible. Instructions addressed to the assistant, sitting inside a message that looks ordinary, in text a person was never going to see.
  • It has a name and a federal taxonomy behind it. NIST calls it indirect prompt injection, and two of the agency’s own worked examples are email: a model in a mail client told to forward messages to an attacker’s inbox, and one told to write itself onward to everybody in your contacts.
  • Filtering is a real defense, and a probabilistic one. The defense that holds whatever the message says is a quarantine, which keeps the message from being read at all. That’s why it’s worth more than it looks.
  • After that it’s settings. An injection travels exactly as far as the assistant was already allowed to go on its own, which makes the autonomy setting a security control rather than a convenience.

What changed when something else started reading first

For you, reading a message and acting on it are two events with a gap between them. A request for a wire transfer lands. You read it, you think, you decide. The gap is where your judgment lives.

For a language model the gap is much narrower. The text it reads and the instructions it follows come down the same pipe. NIST puts the mechanism in one sentence: “Because GenAI models combine the data and instruction channels, attackers can leverage the data channel to affect system operations by manipulating resources with which the system interacts.” (NIST AI 100-2e2025, March 2025, checked September 7, 2026.) Your mail is the data channel. A message is a resource somebody else controls.

Two things follow, and both sit in the same document. They’re what separates this from ordinary phishing.

The first is who does it, and who pays. NIST is precise: “unlike in direct prompt injection attacks, indirect prompt injection attacks are mounted not by the primary user of a model but instead by a third party. In fact, in many cases, it is the primary user of the model who is harmed.” Nobody has to trick you into anything. You’re a bystander here. Somebody sends a message to an address that’s been public for years. The software you pay for reads it as routine, because reading everything is what you bought.

The second is where the instruction sits. You’ve been trained, correctly, to check the sender, hover over the link, and notice the invoice from a domain with one letter changed. Everything you’ve been trained to check is visible. The instruction is in the part of the message that renders as nothing.

There’s a second AI-adjacent attack in your inbox, and the two get confused. A consent screen for an application pretending to be something else is an attack on you, at the moment you connect it. That one belongs with what to check before connecting your inbox. This page is about what happens afterwards, on an ordinary Tuesday, with nothing to approve.

Here’s the part that should steady you, before the rest. Your provider filters before anything else sees the message. Google said on September 1, 2025 that “our protections continue to block more than 99.9% of phishing and malware attempts from reaching users” (Google, checked September 7, 2026), and Microsoft runs its own layer on the same principle. Then read that sentence again. It counts attempts, not your inbox. What survives a filter tuned against billions of samples is the unusual one: low volume, aimed at you, written well. Those are exactly the messages an assistant reads carefully.

What a hostile message is actually asking for

Five things, and only two of them look like theft.

Send the mail somewhere. NIST’s example is the plain one: “a model integrated as part of an email client could be prompted to forward certain emails to an attacker-controlled inbox.” A selection, because taking the lot gets noticed. The thread with the settlement figure in it. Everything from one client. Anything carrying the word “wire.”

Leak without sending anything. This one surprises people, because your sent folder stays empty. NIST notes that “attackers may also be able to exploit features like markdown image rendering to exfiltrate data.” An assistant that renders a message fetches the images in it. A web address can carry text. So an image address built out of your own data carries that data out, quietly, as a request to a server somebody else owns.

Spread. The worm case. NIST writes: “an attacker could send a malicious email that, when read by a model integrated as part of an email client, instructs the model to spread the infection by sending similar malicious emails to everyone in the user’s contact list. In this way, certain malicious prompts could serve as worms.” Your contact list is your clients.

Distort. The quietest one, and the one with no obvious victim. NIST records demonstrated attacks that “cause a GenAI system to produce arbitrarily incorrect summaries of sources, to respond with attacker-specified information, or to suppress or hide certain information sources.” Nothing is stolen. What changes is what you believe about your own inbox: a summary that reads as routine, a message that never surfaces, an answer that’s confidently wrong.

Stall. Instructions that make the model do something slow or refuse a capability. Annoying rather than dangerous, and the least of your problems.

Where the text hides is the part worth carrying around. NIST lists it as a technique in its own right: attackers hide injections “in non-visible portions of a resource,” use multi-stage injections where the first one points at a second, and encode the instructions, “such as in Base64,” then tell the model to decode them. In an email, non-visible means white text on a white background. A font sized to zero. An HTML comment. The alt text on an image. Four hundred lines of quoted trail below a signature nobody scrolls past. All of it is invisible to you. The software reads every line.

This has already happened once in public. On June 11, 2025, Microsoft published CVE-2025-32711, described in its own words as “Ai command injection in M365 Copilot allows an unauthorized attacker to disclose information over a network.” Microsoft rated it Critical at 9.3. The vector Microsoft filed records no privileges and no user interaction required, which is the formal way of saying the victim does nothing. NVD scored the same flaw independently and lower, 7.5 and High, a normal disagreement about scope rather than about whether it was real (NVD, checked September 7, 2026).

Read that as a vendor handling a flaw well. Microsoft recorded the flaw as neither publicly disclosed nor exploited when it published. It fixed the flaw on its own side. Then it published a CVE for a cloud service, under a disclosure practice it adopted voluntarily: “This vulnerability has already been fully mitigated by Microsoft. There is no action for users of this service to take. The purpose of this CVE is to provide further transparency.” (Microsoft, checked September 7, 2026.) The lesson is the arrival route. A large, well-resourced vendor shipped an AI assistant wired into its customers’ mail. An email could make that assistant give things up, and the customer’s own caution never got a chance to weigh it.

Why asking the assistant whether a message is safe is the wrong move

The instinct is immediate, and almost everybody has it. Something looks off, so you ask the thing that reads your mail: is this real?

It fails for two separate reasons, and the second is the serious one.

The first is that a summary is the wrong instrument. What settles a message’s identity sits in the material a summary throws away. The sending address, rather than the display name. A reply-to that differs from the from. A domain with a letter inserted. Where the link goes, rather than what it says. Whether the formatting matches the eleven previous invoices from that vendor. A summary of a fraudulent invoice is a competent, accurate description of an invoice. It answers the question you asked, and the question you asked was the wrong one.

The second runs deeper. If the message is hostile, you’re asking a witness that’s already been got at. The distortion attack above exists to shape what the model says about the source. An assistant that has read a message full of instructions can hand back the assessment those instructions asked for. That assessment reads exactly like every other one you’ve had.

So here’s the rule, flatly. You judge a message on the raw message. You confirm it somewhere other than email. The FBI’s guidance for the family of fraud this feeds is one line long: “Use secondary channels and/or two-factor authentication to verify requests for changes in account information.” (FBI IC3, September 11, 2024, checked September 7, 2026.) A secondary channel means a phone number you already had, from a record that predates the message. The number in the signature arrived with the message. So did the reply address, which reaches whoever controls the thread.

The scale is measured. In its 2025 annual report the IC3 recorded 191,561 phishing and spoofing complaints, the most of any crime type by count. It recorded 24,768 business email compromise complaints carrying $3,046,598,558 in losses, second only to investment fraud. Complaints flagged as AI-related came to 22,364, with adjusted losses of $893,346,472 (FBI IC3, checked September 7, 2026). This is ordinary business crime with a large number attached. Small professional firms sit in the middle of it, because they hold other people’s money in motion.

The defense that does not depend on noticing

There are two families of answer, and one is stronger than the other.

Catch it. Classifiers read incoming text and score how much it looks like an injection. Models are hardened against following instructions found in data. Sanitizers strip risky formatting. Google publishes its own version, in more detail than most vendors give (Google Workspace, checked September 7, 2026):

  • content classifiers
  • security thought reinforcement
  • markdown sanitization with suspicious URL redaction
  • a user confirmation framework
  • end-user notifications
  • model resilience

Google’s own consumer wording is honest about the shape of that. Gemini in Workspace “may filter or block some responses if malicious activity is detected.” Ask Gemini to summarize messages and one of them is malicious, and Gemini “may not respond to your prompt for safety reasons” (Google, checked September 7, 2026). That’s a genuine defense doing real work, and every word of it is probabilistic. Note the timing, too. The phrase “may not respond” describes a defense that fires after the message has been read.

Don’t read it. The other family sorts on something duller and much harder to fake: whether you have any history with this sender. A message from somebody new is held before anything processes it. The holding comes first, ahead of the summary, the score, the filing and the model. The message sits outside, and you decide whether to let it in.

The difference shows up in the attacker’s job. Against a classifier, the job is to write something the classifier scores as clean. That’s an ordinary contest, and both sides keep getting better at it. Against a hold there’s nothing to score, because nothing was read. The attacker’s problem stops being phrasing. It becomes getting a relationship with you first.

NIST lands in the same place, from the engineering side. Its mitigation section lists the training, detection and filtering approaches. Then it says what they add up to: “Because current mitigations do not offer full protection against all attacker techniques, application designers may design systems with the assumption that prompt injection attacks are possible if a model is exposed to untrusted input sources.” A quarantine is that assumption turned into a piece of the product. Keep untrusted input away from the model, and whether the model can be fooled stops mattering.

The cost is real, and here it is. A genuine first-time client is a stranger too. The prospect who found you through a referral, the opposing counsel writing for the first time, the new vendor sending an onboarding pack: all held, all waiting. For most small professional firms that’s a good trade. Almost all of your daily mail comes from people you already have history with, and the held list behaves like an intake queue rather than a discard pile.

Two neighbors sit close to it, and the names invite the confusion. Your junk folder answers a different question. Junk filtering asks whether a message is unwanted. A hold asks whether the software has any reason to trust the sender. Your filtering service also runs an administrator quarantine at the platform level, which stops mail before your mailbox ever sees it. Most people meet that one only when something goes missing. What a quarantine folder is covers it, including where it actually lives and who can release from it.

How far an injection can travel is a setting you already have

Suppose one gets through anyway. One question decides how bad the day is, and it’s a question about your settings rather than about the message.

An injection does exactly what the assistant was already permitted to do on its own, and stops there. Put replying and sending at a level where the software prepares and waits for you, and the worst case is a draft that looks slightly odd sitting in a queue. Put them at a level where the software acts on its own, and the worst case is a message already sitting on somebody else’s server, out of your reach.

OWASP reaches the same conclusion in its catalog of these attacks. Among the mitigations for LLM01:2025 it lists both “Enforce privilege control and least privilege access” and “Require human approval for high-risk actions” (OWASP, checked September 7, 2026). Google’s layered strategy includes a user confirmation framework for the same reason. The industry’s answer to a problem it can only partly solve is to shorten the leash on the actions that matter.

The practical version draws a line in a specific place. Actions that stay inside your account can run further ahead than actions that leave it. Turning everything down throws away the reason you bought the software.

Filing, labeling, snoozing, summarizing, drafting, extracting a task. These are recoverable. A wrong one costs you a few seconds and leaves a trace.

Sending, replying, forwarding, sharing, anything that changes a payment instruction. These leave the building. A wrong one is somebody else’s copy now.

Keep the second group under review for longer than the first. That’s the whole of the advice, and it earns its keep on plain accuracy grounds even if injection never crosses your mind. What each level actually means is set out in setting how much your inbox does on its own. The line between what a record can show you and what a reversal can actually take back is what an activity log lets you check, and what undo can reverse. One half of that line matters here. A reversal reaches what the software did inside your account, and it stops at the edge of your account. A message sitting on somebody else’s server stays there.

What to do with the one that got through

Six steps, in this order. The first one is the one people get wrong.

  1. Don’t hand it to any AI tool. Not the assistant, not a chat window, not a browser extension that offers to check it. Forwarding a hostile message into an AI tool to ask about it delivers the payload by hand, and a forward carries the hidden parts along with the visible ones. The same caution covers the attachment, which is a document a stranger composed and which asking a question about an attachment treats at length.
  2. Look at the raw message. The full sending address, the reply-to, the domain, where the links actually point. If you’re comfortable viewing the message source, the injected text is usually sitting there in plain sight, a fast way to stop wondering.
  3. Verify somewhere other than this thread. Call a number you already had. If the message asks for anything to change about how money moves, that call is required. It’s the oldest control against this family of fraud, and still the one that works.
  4. Report it where it does some good. Both Gmail and Outlook carry a report control in the message itself, and using it feeds the filter that protects everyone on that platform, including you next week. If money was requested or moved, file at ic3.gov. The FBI’s Recovery Asset Team initiated 3,900 Financial Fraud Kill Chain actions in 2025, against $1,163,919,846 in attempted theft, and froze $679,013,183 of it, a 58% success rate. Its own advice on the timing is blunt: “If you discover a fraudulent transfer, time is of the essence. Immediately, contact your financial institution and request a recall of the funds along with any necessary indemnification documents.”
  5. Leave it where it is. If your assistant held it, releasing it to take a closer look defeats the mechanism you paid for. A tool that does this properly gives you a way to read the held message without processing it. If yours works differently, open it in your provider’s own web interface rather than releasing it.
  6. Tell whoever it impersonated. If a message claimed to be your client’s bookkeeper, that bookkeeper’s mailbox may be the one that is compromised. You may be the second firm it reached rather than the first.

What Point holds at the door, and what that costs you

Point is an AI email client, so everything above applies to Point too, and the design decision that matters here runs in the same direction this page argues for.

  • A sender with no history is held at the door rather than read into anything, which is how a stranger is kept from instructing your inbox. The wording is deliberate. The holding happens before the reading, which removes the attacker’s phrasing problem rather than making it harder.
  • Mail that carries the marks of a manipulation attempt is set aside too, for the same reason and by the same route. It becomes something you decide about rather than something you react to. Each held message carries the reason on its card, and injection risk is one of the reasons it can carry.
  • You can read a held message without anything processing it. Review opens a sealed view where nothing runs. The raw text is visible to you, and any hidden instructions are visible with it. That solves the problem step two of the last section describes. Looking closely usually means handing the message to exactly the reader you were protecting.
  • Release waits for you. Releasing a message returns a false positive to normal processing. Whatever you leave in the hold stays outside.
  • How far Point goes is set separately for each kind of action, and each one starts on review. Out of the box Point prepares and waits. You promote the kinds of work you’re ready to hand over, one at a time, which means the reach described two sections up is something you set rather than something you inherit. Approving anything new before it acts is where that starts, and the activity log behind every action is where you go to see what happened.
  • The cost is a first-time client waiting at the door. For a firm whose ordinary correspondents all have history with you, that’s a queue to glance at once a day. For a business whose mail is mostly inbound strangers, it’s a heavier tax, and worth knowing before you decide rather than after.
  • The outer edge is still the grant you issued at your provider. Point works inside the permissions on that consent screen, which is why the screen is worth two minutes of reading, and where your mail goes when AI reads it picks up on the other side of it.

Everything Point does with a connected mailbox is listed on the benefits page.

Common questions

Is prompt injection a real risk for a seven-person firm, or a research problem?

Both, and the research came first. NIST documents the attack class in a published taxonomy, OWASP catalogs it, and it has produced at least one critical CVE in a shipping product from Microsoft. What makes it a small-firm problem rather than an enterprise one is arithmetic. A professional services firm holds client money in motion, keeps a public address, and has nobody whose job is watching for this. The FBI recorded $3,046,598,558 in business email compromise losses in 2025 across 24,768 complaints, which is what the underlying fraud costs before AI is involved on either side of it.

Can I forward a suspicious email to my assistant and ask whether it is a scam?

That’s the one move to avoid. A forward carries the invisible parts of the message along with the visible ones, so you’re delivering the payload to the reader by hand and then asking that reader for an opinion. The answer you get can be the answer the message asked for, since NIST records demonstrated attacks that make a model produce arbitrarily incorrect summaries or respond with attacker-specified information. Judge the message on the raw text. Confirm anything that matters on a phone call to a number you already had. Use your provider’s report control.

My spam filter already catches this. Why is anything else needed?

It catches most of it, and most is the operative word. Google’s own figure is more than 99.9% of phishing and malware attempts blocked, and that’s a share of attempts rather than of your inbox. What’s left is the targeted, low-volume, carefully written mail that a filter tuned on mass campaigns is least likely to recognize. That’s also exactly the profile of a message written to instruct an assistant rather than to fool a person. Your spam filter is answering a different question. It asks whether a message is junk. A hold asks whether the software about to read it has any reason to trust the sender.

What happens when a real new client writes for the first time?

They wait. That’s the honest cost of the mechanism, and the reason the queue should be visible rather than silent. A held list is worth checking once or twice a day, in the same way a voicemail light is, and a genuine first message from a prospect is easy to recognize when you see it. Release the sender once, and the history exists, so later messages arrive normally. If your business runs on cold inbound rather than on an established book of clients, weigh this more heavily before you buy. The queue is your front door rather than an edge case.

I have had an AI assistant connected for months. Should I go back and check something?

Two things, and both are quick. The first is a look in two places. Check the activity log for actions you didn’t expect, particularly forwards, sends and rule changes. Then check your mailbox’s own forwarding and filter settings at the provider, which sit outside any assistant’s log and are a classic destination for a successful attack. The second is your autonomy settings. Look at where they actually sit today rather than where you remember putting them, because they tend to drift upward as trust builds, and the settings that let an assistant act unattended are the ones that decide the worst case.

The short version

  • A message can be written for the software that reads your inbox rather than for you. NIST calls it indirect prompt injection, and its own examples are a mail assistant told to forward your messages to an attacker or to write itself onward to your contacts.
  • The instructions sit in the parts of a message that render as nothing: white text, zero-height elements, alt text, a buried quoted trail. You were taught to check what’s visible, and this lives past the edge of what renders.
  • It has already happened in public. Microsoft published a critical CVE in June 2025 for an AI command injection in M365 Copilot that required no user interaction, fixed it on its own side, and said so.
  • Never ask the assistant whether a message is legitimate. A summary discards the evidence that decides it, and a hostile message may have already shaped the answer you get back.
  • Detection is a real defense, and a probabilistic one. Holding a stranger’s message before anything reads it is a structural one, and NIST’s own conclusion is to design as though injection is possible whenever a model sees untrusted input.
  • How far a successful injection travels is a setting you already own. Keep the actions that leave your account under review for longer than the ones that stay inside it, and the worst case stays a draft rather than a sent message.
  • Verify anything about money on a phone call, to a number you had before the message arrived. File at ic3.gov if money was requested or moved.

Maybe the bigger question underneath this one is whether to connect an AI tool to a mailbox full of client business at all. The whole list is worked through in the security questions worth settling first. And why you are right to be careful with AI starts from the feeling rather than the checklist.

Ready for a calmer inbox?

Join the private beta

We're onboarding a few teams at a time. Leave your email, confirm it once, and we'll send an invitation the moment a place opens.

By joining you agree to our privacy policy.

Private beta

What you're joining

It runs on the mail you have

Point sits on top of Gmail or Outlook. Your address, your history and your contacts stay exactly as they are, so there is nothing to migrate.

You set how much Point does

Out of the box everything waits for your review, replies included. You hand over only what you trust, one kind of work at a time.

Join the private beta

We're onboarding a few teams at a time. Leave your email, confirm it once, and we'll send an invitation the moment a place opens.

By joining you agree to our privacy policy.