Skip to content

Lessons from the beta, and the places we had it wrong

On this page

More firms are being admitted to Point’s private beta between now and the end of the year, so this is the account of what the first ones found. It includes the places we had it wrong, because those are the parts an incoming firm can actually use. There are no percentages in it, and the reason for that is the first section.

  • A beta this size cannot produce an honest statistic. What it produces reliably is the second time something breaks the same way at an unrelated practice, and nothing is below unless it reached that bar.
  • We had prepared for the wrong objection. Almost nobody led with a question about models. They asked who at our company can read a thread, and whether the engagement letter has to change.
  • Week one undersells itself and we did not warn anybody. On the first morning Point has read your mailbox but has not yet watched your practice, and a firm not told that reads two ordinary weeks as a verdict.
  • A protective feature is judged on the one real message it held, never on the ten hostile ones it caught, because the catches are invisible and the mistake has a name on it.
  • Almost nothing moved off its starting setting, and not because firms had decided against it. Nothing ever put the decision in front of them.
  • The most frequently requested feature was a different product.

Why there are no numbers here

The firms inside a private beta are not a sample. They are the practices that said yes first, which selects for curiosity, spare attention and an owner who already suspected their mailbox was costing them something. A percentage drawn from that group would describe the group rather than the software, and publishing one would be the most flattering thing on this page and the least true.

The other reason is seasonal. No firm has yet carried Point through a February. The beta has run through a summer and into a fall, so anything said here about how this behaves on March 11 would be a prediction wearing the clothes of a finding. The seasonal reasoning that follows from that, including why the fall is the cheap month to start in, belongs to getting ready for the 2027 tax season.

What a beta of this size does produce, and produces well, is repetition. One firm’s problem is an anecdote about that firm. The same problem at a practice with different clients, a different provider and a different owner is a design fault. Everything below happened more than once, at practices with nothing in common except the kind of work they do.

The objection we were ready for

We had built a security answer for a question about models. Which model, hosted where, does it train on our mail. Those are reasonable questions and the answers were public before the first firm connected, in privacy and subprocessors.

Almost nobody led with them. What came instead, in some order, in nearly every first conversation:

Who at your company can read one of my client threads, and in what circumstances. If a client asks me whether AI touched their return, what am I supposed to say. Does my engagement letter have to change, and by when.

Two of those are about people and the third is about a document with a December deadline on it. None is a model question. And the third one is the one that actually stops a firm, because it is the only one with a date attached: the letters are signed in December, and from that point the wording governs every engagement of the season. That timing, and what a letter can and cannot carry, is what to put in your engagement letter about AI.

We had answers for all three. They were in a legal document, in the vocabulary of a legal document. Expecting a practice owner to work through a privacy policy at nine at night is asking them to do our job for us, so we wrote the answer out in the language a firm actually uses, which is how we think about data. It points at the specific paragraph behind every claim it makes, including the ones that are narrower than the slogan version.

The finding generalizes past us. The question that decides whether a firm can use a tool in this category is not a technical question, and a vendor whose only answer lives inside a contract has not answered it. The rubric worth applying to us and to everyone else is questions to ask any AI tool about your data.

One more thing surprised us here. The firms most exercised about client data were not the least technical ones. They were the ones where something had already gone wrong, usually a near miss involving a fake invoice or a request that turned out not to come from the client. A near miss makes the question specific, and a specific question is far easier to answer well than a general unease.

What we should have said sooner

The most common disappointment in the first two weeks was not a bug. It was that the first morning was less impressive than the demo, and nobody had been prepared for why.

On day one Point has read your mailbox. It has not yet watched your practice. The ordering you get is competent and general: it can tell a person from a payroll bureau, a question from a receipt, a request from an acknowledgment. What it cannot yet know is that one particular client’s attorney outranks nearly everything, or that a certain sender writes urgently about matters that can comfortably wait until Thursday. That comes out of weeks of ordinary traffic, and there is no settings page that shortens the wait.

A firm told this in advance treats week one as calibration. A firm not told reads the same two weeks as evidence, and the second reading is fatal, because the firm is right about what it saw and wrong only about when to judge it. That was our error rather than theirs, and it is now said out loud before a mailbox is connected. Here is the version we say.

The first two weeks are for correcting, not for evaluating. The corrections that pay are small and specific: fixing a summary that missed the point of the thread, pulling up something that was ranked too low, stating once in plain words that a certain kind of message always matters. Each of those is read as a standing instruction rather than as a one-time fix, which is when a summary learns from the corrections you make.

The corrections that do not pay are the general ones. “Be smarter about clients” is not a correction, and neither is a rating out of five. Neither one tells anybody which thread went wrong.

And a firm that connects a mailbox and then touches nothing for a month has, at the end of that month, exactly the general version it started with. That was the single most reliable predictor of a firm going quiet, and it is not a subtle one. The day by day account of what the early days are actually made of is day one, connecting the inbox you already have, and the mechanics of joining and the first Monday belong to opening up to more firms.

The shape of a practice’s mail

We designed against a persona and then met the mail. Most of what we had was right. Four things were not visible from outside a practice, and all four turned up at more than one firm.

Most of the mail about a client is not from the client. One engagement generates traffic from the client’s bookkeeper, their payroll bureau, their attorney, a broker, a bank, and a portal sending notifications on their behalf. Six addresses, one matter, and only one of them carries the client’s name. Software that reasons about senders gets this wrong every time, because the unit that matters in a firm is the matter rather than the correspondent. This is the part of the ranking that most needed to learn a specific practice rather than practices in general.

Important and urgent come apart harder here than we expected. The genuinely urgent mail in a firm mailbox is frequently a machine with a deadline behind it. The genuinely important mail is frequently a client being polite. A message that opens “no rush at all, whenever you get a minute” is, in a professional-services firm, often the most expensive item in the pile, and the ordinary inbox rewards it with the worst position. Untangling those two axes is important versus urgent.

The request is rarely in the first paragraph. Real client mail buries the ask under context. Paragraph four, after two paragraphs about the summer and one about a form they are fairly sure they already sent. A practice mailbox delivers a dozen of those a day and labels none of them, which is why lifting the ask with the date the client actually named is one of the four jobs the beta holds itself to rather than a nice touch.

Silence carries more weight than we had allowed for. The thread that never came back is the one that surfaces in April, and in a practice the cost of a missed reply is asymmetric: a client who has gone quiet on a document request is a deadline problem three months before anybody notices it is one. Threads that stop, and what to do about them, are unanswered emails and the threads that come back. What the firm mailbox ends up being the only record of is running the firm inbox.

All four point in the same direction, and the direction they point is the depth this quarter is being spent on rather than a longer feature list. That plan is what we’re building this quarter.

What a wrongly held message costs

Mail from a sender Point does not recognize is set aside before anything reads it for meaning. The reason is not spam volume. It is that a stranger should not be able to issue instructions to your inbox simply by writing to you.

The finding here is about arithmetic rather than about the mechanism. A protective feature that catches nine hostile messages and wrongly holds one real client is not experienced as nine to one. It is experienced as one. The nine were invisible, and the one had a name on it and a phone call after it. That reaction is not irrational, either. A held invoice from a fake vendor costs a firm nothing it will ever learn about. A held message from a real client costs a returned call and a small dent in the thing a practice actually sells, which is the sense that you are on top of their affairs.

Two things follow. The first is ours: a feature like this has to be tuned against its false positives rather than its catch rate, and the pile it holds has to be somewhere a person genuinely looks rather than a folder everybody assumes is empty.

The second is yours, and it takes about five minutes in a first week. Put the people you cannot afford to have held in front of the door before the door gets a chance to hold them. The clients who are mid-engagement, the two firms you refer work to, your own bank, your own software vendors. That list is VIP mail that never gets buried. Then look at the held pile daily for the first two weeks, not because it will be full but because those are the two weeks in which you find out what your firm’s version of a stranger looks like.

Whatever was held, released or acted on is a timestamped line in the activity log, and most of what Point did can be put back. The exception is delivery. Once a message has landed on somebody else’s server it is out of reach, and no product in this category recalls it.

Nothing moves off review by itself

How far Point goes is a separate setting per kind of work, and every one of them ships on review, which means Point prepares the thing and then waits. Pushed to the top of its range, a setting stops checking in before it acts.

The finding is that months in, most settings sat exactly where they had started. Not because firms had weighed the decision and declined it. Because nothing ever put the decision in front of them. A settings page turns out to be a poor place to earn trust: it asks a busy person for an abstract judgment about something that has not happened yet, in a room with no evidence in it.

What moved a setting, on the occasions one moved, was almost always the same sequence. Somebody read a week of the activity log, noticed they would not have done anything differently in any of it, and raised exactly one setting the following morning. That is the whole mechanism, and it suggests an order worth copying.

Raise the routine ones first, the ones whose worst case is untidiness rather than embarrassment: filing, ranking, and an acknowledgment you have already worded yourself. Leave the settings that write to a client where they are for longer than feels necessary, and move one only after reading a stretch of drafts you would have sent unchanged. Raise one at a time, because the point of a per-action setting is that a bad week for one of them is not a bad week for all of them. The full anatomy of the dial is setting how much your inbox does on its own, and the narrower question of what has to be approved before it happens at all is approving anything new before it acts.

The practical version of all this is a standing ten minutes on a Friday to read the week’s log and decide whether anything moves. It converts a pending decision into a recurring one, and pending is where these decisions go to die.

The request we kept turning down

The most frequently requested thing was not a feature. It was a different product.

Some version of: show me, per client, which of the eleven documents we asked for have actually arrived. Or: show me how much work is stacked in front of each preparer, so I know in February which returns are not going to be finished. Both are real problems, and both are among the more expensive things a small practice carries around unmeasured. Neither is an email client’s job, and no amount of making a mailbox cleverer turns it into a register or a workload board. Why a document register belongs on one list rather than in the mail is one PBC list, not five copies of it in the mail, and the rest of what Point is deliberately not is set out in what we’re building this quarter.

The lesson we took from it is about timing rather than scope. A firm that works out in month two that it had been shopping in a different category has spent the cheapest month of its year finding that out. So the boundary now belongs in the first conversation, said plainly, even when saying it costs the conversation. It costs fewer of them than you would expect, and the ones it costs were going to end the same way in March.

What separated the firms it took with

Three differences, and none of them is firm size or how technical anybody was.

Somebody owned it. In the firms where it stuck, one named person did the correcting, checked the held pile and read the log. Not a project and not a rollout. A habit, held by a person, for about ten minutes a day for two weeks. Where it was everybody’s job it was nobody’s, and the mailbox kept the general version described above.

They started with slack in the calendar. A firm with room treats a wrong ranking as information. The identical event during a compressed week reads as an interruption, and an interruption gets a tool closed rather than corrected. That is the whole seasonal argument, and the version with the dates in it is getting ready for the 2027 tax season.

They had already had the client-data conversation. Not resolved elegantly, just had. A firm still carrying that question unopened tends to use the software tentatively, which produces a thin result, which then gets read as the product being thin. Two separate questions, and the order matters: settle the first, then evaluate the second. Is it safe to use AI with client financial data is where that starts.

For the record, here is what predicted nothing. Firm size, inside the range Point is built for. Whether the firm was on Google or on Microsoft. How comfortable anybody said they were with AI in general. And whether the owner described themselves as technical, which was the most consistently useless signal we had.

If you are weighing this for your own practice, the terms, the fit and the way in are opening up to more firms, the full capability list is on the benefits page, the practice-shaped version is Point for accountants, and what actually changes when a mailbox is connected, which is less than most firms expect, is how to switch to an AI email client. Point is a private beta, priced per seat, and its waitlist is the only way in.

Common questions

How many firms are in the private beta?

Not a number we publish, and this post carries none deliberately. What is worth knowing about the size is what it does to the evidence rather than the figure itself: the beta is small enough that any single firm’s experience is an anecdote, which is why nothing above is here unless it happened at more than one unrelated practice. Who is being admitted and on what terms is opening up to more firms.

Has anyone run a full tax season on Point?

No. The beta has not yet been through a February, so anything anyone tells you about how this behaves in the last week of March is a prediction rather than a finding, ours included. That is precisely why a firm joining now is pointed at December: the aim is to be settled before a season rather than to be evaluating something during one.

What did you actually change because of all this?

Most of it changed what we say rather than what we ship. We wrote our data answer down in the language a practice uses instead of leaving it in a contract. We tell firms what week one really looks like before a mailbox is connected. We name the boundary in the first conversation rather than letting it be discovered. And we concentrated this fall’s recruiting into November so an incoming firm still has a quiet December to be a beginner in. The product work these findings point at is scoped in what we’re building this quarter, and it is depth on four jobs rather than a fifth one.

Should we wait until the beta is further along?

Depends what you would be waiting for. If it is a published price or a certification report, waiting is the right call, because neither exists today and no badge is going up here before there is a held certification behind it. If it is for the software to know your practice, waiting does not help, since that only begins when your mail starts going through it. And if what is unfinished is your firm’s own conversation about software reading client mail, then that conversation is the work, and it is a separate piece of work from this decision.

We are not an accounting firm. Does any of this apply?

Most of it. The findings about week one, the held pile and the settings are about how people adopt software rather than about tax. The section that is genuinely practice-specific is the shape of the mail, and even there the underlying pattern holds anywhere a single matter generates traffic from six different addresses. This season’s beta concentrates on accounting and tax firms because one kind of year gives a clear read, and the rest of who Point is built for is on who it’s for.

What is the most common reason a firm stops?

Nobody owning it. A mailbox gets connected, nothing gets corrected for a month, the ordering stays general, and the firm draws a reasonable conclusion from what it saw. That failure is solvable, and the solution is ten minutes a day from one named person for two weeks. The second most common reason is not solvable by us, and it is the firm having been shopping for a document register or a workload board, where an email client is the wrong shape whatever it does.

Do you take feedback from beta firms, and what kind is useful?

Yes, and the useful kind is narrow. The message that got ranked below something trivial, with the thread still open in front of you. The ask in paragraph four that never became a task. The sender who was held and should not have been. General praise and general frustration are equally hard to act on, and a rating tells us nothing we can fix. A single thread we can open is worth more than a page of impressions.

The short version

The first firms in Point’s private beta taught us more about ourselves than about the software. We had prepared for a question about models and got three questions about people and a document, one of which carries a December deadline. We let firms walk into a first week that undersells itself without telling them why, which is our error and is now said out loud before anyone connects: weeks one and two are for correcting, and a mailbox nobody corrects stays general forever. Practice mail has a shape we had only partly seen, where most of the traffic about a client comes from somebody other than the client, the urgent message is often a machine and the important one is often a client being polite. A feature that protects you is judged on the one real message it wrongly held, so put your VIPs in front of the door and read the held pile daily for two weeks. Almost no settings moved on their own, because a settings page is a poor place to earn trust and a week of the activity log is a good one. The most requested feature was a different product, and saying so early costs less than letting a firm find out in month two. And the firms it took with were not the technical ones or the large ones. They were the ones where a single person spent ten minutes a day for two weeks, in a month that had room in it.

Ready for a calmer inbox?

Join the private beta

We're onboarding a few teams at a time. Leave your email, confirm it once, and we'll send an invitation the moment a place opens.

By joining you agree to our privacy policy.

Private beta

What you're joining

It runs on the mail you have

Point sits on top of Gmail or Outlook. Your address, your history and your contacts stay exactly as they are, so there is nothing to migrate.

You set how much Point does

Out of the box everything waits for your review, replies included. You hand over only what you trust, one kind of work at a time.

Join the private beta

We're onboarding a few teams at a time. Leave your email, confirm it once, and we'll send an invitation the moment a place opens.

By joining you agree to our privacy policy.