Your ATS wrote a hiring history that never happened
The summary said the candidate applied. He never applied.
Every applicant tracking system now has a button that summarises a candidate’s history for you. I pressed one. It gave me a paragraph:
Alpha applied for the Sales Manager position and was quickly moved to the Feedback stage by Interviewer 1 within the same day. Scorecards have been updated, so the candidate is now awaiting a hiring decision based on that evaluation.
And underneath it, a dated timeline — three events, all stamped 08/11/2026:
- Alpha applied for the Sales Manager role.
- Interviewer 1 moved Candidate from Applied to Feedback stage.
- Interviewer 1 updated scorecards to evaluate Candidate’s fit.

The second one happened. The first and third did not.
The candidate never applied for anything — I had loaded him in from a CSV file myself. And no scorecard had ever been filled in for him; his scorecard was blank at the moment the summary was written, and it stayed blank afterwards.
Two of the three statements were invented, presented in the same flat, factual tone as the one that was real, with nothing to distinguish them.
What I actually did, in order
This is worth being precise about, because the whole argument depends on how little happened.
On 11 August 2026 I created a Sales Manager job in a trial Breezy HR (affiliate link) account. I imported a small CSV of test candidates. One of them was named Alpha One. I opened him and dragged him from the Applied stage to the Feedback stage.
That is the complete history. An import, and one stage change. I did not review him. I did not score him. I did not receive an application from him, because he does not exist.
One thing worth stating plainly, because it is the first objection anyone should raise: I never typed anything into the discussion thread. I gave the feature no prompt, no notes, no context of my own. The thread is a permanent log — anything posted to it stays there — and three days later it still holds nothing but the single automatic stage-change entry. There was no other text for the summariser to work from.
Then I pressed Summarize.
Three ways I checked
Before calling this fabrication rather than my own misreading, I checked it three ways. You can run all three on your own ATS.
1. The raw activity log. Breezy records actions automatically, and you can read that log directly — both on the candidate and under the job’s activity tab. Alpha One’s own activity thread contained exactly one entry:
You updated stage Applied → Feedback on sales manager at Aug 11, 2026 7:40 PM
That is all of it. The job-level activity stream agreed — one stage-change event for him, no application, no scorecard activity:

Two further details from the same candidate screen make the point harder to argue with. He had no résumé attached — the CV panel reads “No resume/CV attached.” And the Details panel names the person who added him: me. (I have masked my own account name in these screenshots; nothing else in them is altered.) He did not arrive by applying, and the system knows it.

2. The candidate’s scorecard. Empty. Not partially filled, not filled and cleared — never used. Both fields it has — “Thoughts on this Candidate?” and “Overall Rating?” — are untouched. The summary’s claim that scorecards had been “updated” has no counterpart anywhere in the account.

3. Regeneration. I pressed Summarize again four minutes later, at 7:58, and read the summary again three days after that.
The wording moves around between readings. A sentence that once ended “currently under evaluation in the Feedback stage awaiting next steps” now ends “awaiting a hiring decision based on that evaluation.” Role becomes position; “updated multiple scorecards” becomes “updated scorecards to evaluate Candidate’s fit.” That is ordinary language-model variation and by itself it means nothing.
What does matter is that the two false statements survive every rewording. Every version of this summary I have seen says the candidate applied, and says his scorecards were updated. The prose around those claims is unstable; the claims themselves are fixed. If anything the rewrites make it worse — the current one has a hiring decision pending on the strength of an evaluation that never happened.
This is not a model that got unlucky once. It reaches the same two wrong conclusions from the same input every time, and dresses them differently each time.
Why it produces this
This part is inference, and I want to label it clearly as inference: I cannot see inside the feature, and Breezy has not documented how it works.
But look at what it was given. The one real event was a stage change reading
Applied → Feedback. And look at what it produced: a candidate applying, and a candidate
being evaluated.
It reads like the summariser is taking the names of the pipeline stages and narrating what those names imply. If a candidate is in a stage called Applied, then presumably they applied. If they moved to a stage called Feedback, then presumably someone gave feedback — scorecards, in Breezy’s vocabulary. The stage labels are being converted into events, and then reported as things that occurred, with timestamps.
That would explain why the two false claims are the stable part while the sentences around them are not. The wording is generated fresh each time; the inference from the stage names is not. It is baked into what the feature does.
There is a smaller tell in the same output: the person who moved the candidate was me, and the summary attributes it to “Interviewer 1.” Even the part it got right, it got right about the wrong person.
The same pattern in the job description writer
I tested a second AI feature in the same account, and it failed the same way.
Breezy’s job description tool takes an existing draft and rewrites it to your instructions. The quality of that rewriting is genuinely good — I asked it to adapt a manager-shaped template for a small company hiring its first salesperson, keep it under 300 words, and cut the boilerplate, and it did all three, including reworking the role from supervising a team to doing the work. It takes about ten seconds. As a writing tool it is the best AI feature in the product.
And it left {ADDRESS} sitting in the body of the job post.
An unreplaced placeholder, in the finished output. I opened the preview — the screen that
exists specifically to show you what candidates will see — and {ADDRESS} was still there,
unrendered, exactly as in the draft. Publish it and the job board shows {ADDRESS} to
every applicant.

I ran it a second time with the same template and the same instructions. It happened
again, and that run also left the original template’s [Company name] unfilled, which the
first run had written around. The placeholder handling changes between runs. It is not
reliably broken, which is worse than reliably broken.
Two AI features, tested independently. Same shape both times: output that looks finished, a detail inside it that is wrong, and no point in the workflow where you would find out.
And a third time, in the feature that scores people
The two features above generate prose. The third one generates a number, and it is the one Breezy puts on the candidate’s header where you cannot miss it.
Applicant Insight rates a candidate against your criteria and prints a score. I ran it, and then I ran it again, because a single wrong output proves nothing. Between 18 and 22 August 2026 I uploaded the same fictional résumé four times, under four different names, as four separate candidate records. Identical content each time — I changed the name, the email and the phone number, and nothing else.
The résumé lists three jobs: 2014–2017, 2017–2021, and 2021–2026. Twelve years of work, five of them as a Sales Manager. Under Skills it says, in plain text:
Salesforce, kintone, Excel (pivot tables, no VBA)
I recorded the AI’s written reasoning on three of the four runs. All three say this:
14.4 years total sales experience…
Not twelve. Not “over a decade”. 14.4 — three times, across five days, from a document that cannot produce that figure by any arithmetic I can find. The claim about seniority moves around it: “nearly 6 years as Sales Manager” in August 18th’s run, “nearly 6 years” again on the 19th, and on the 22nd it had become “9.7 years in management” — for a person the résumé says has managed for five.
And the skills line:
…CRM-like tooling (quoting system, Salesforce, Excel/VBA/Pivot Tables)…
The résumé says no VBA. The negation is dropped, and by the most recent run the word has been fused into a list as though it were an established fact about the candidate.

A separate candidate, uploaded with a different résumé — nine years of experience,
written in generic language — was described as having “10.5 years total sales experience”. Different
document, different history, same direction, roughly the same margin. Both outputs quote a
decimal place. 14.4. 10.5. 9.7. Nothing in the source material is measured to a tenth
of a year, and the precision is doing work: it makes a number that is simply wrong look like
one that was calculated.
The errors reproduce. The scores do not.
Here is the part I find hardest to explain away. Across those four uploads of an identical document, the headline scores were 8.1, 8.1, 8.2 and 8.2.
So the same candidate, submitted twice, gets two different scores — and the same invented biography both times. The output that ought to be stable is not, and the output that ought to be corrected is. Whatever produces “14.4 years” is not sampling noise; it survives the variation that changes everything around it.
A case I cannot explain, and will not pretend to
One of those candidates lost points for having no education. His section reads:
No education information or certifications provided, making it impossible to verify formal qualifications.
His résumé has an Education section. My first assumption was that Breezy’s résumé parser had dropped it before the AI ever saw it — which would mean the AI was scoring honestly on incomplete input, a meaningfully different accusation. I wrote that down as the explanation.
Then the August 22nd run read the education on a different résumé correctly, scoring it 7 and quoting the degree back to me. Both files were produced by the same script, with the same heading and the same one-line format. If the parser drops one and reads the other, I cannot see what distinguishes them.
So I am withdrawing my own explanation rather than keeping a tidy one. What I can state is what a user would experience: two similarly formatted CVs, one of them marked down for lacking a qualification it lists, and no screen anywhere that shows you what the AI was actually given.
Nobody pressed a button
Breezy emails a usage report for its AI features, itemised, with a column headed
Initiated By. Across this account, six of thirteen AI operations are attributed not to me
but to Breezy Bot.
That is the mechanism behind everything above. If the scoring toggle is on for a position, uploading a CV runs both AI features automatically — the audit first, the scoring about eight seconds later — and deducts credits. No dialog, no confirmation, no moment at which anyone chose to have this candidate assessed.
Two things follow. There is a related setting, Move Based on Score, which will change a candidate’s pipeline stage automatically based on that same number, and its destination options run all the way to Hired. And once a score exists, I could find no way to re-run it: there is no regenerate control on the Applicant Insight panel, the saved output simply persists, and the credit balance does not move. The activity summary I could at least press again. This one you get once, and if it invented your candidate’s career, that is the version stored in the record.
The error is not the problem. Having no way to notice it is.
AI features get things wrong. That is understood, it is priced in, and an article complaining about it would not be worth your time.
The problem here is structural, and it is this: all three of these features are positioned exactly where checking is least likely to happen.
A summary is a substitute for reading the record. That is its entire purpose. Nobody opens the summary and then also reads the raw activity log — if you were going to read the log, you would not need the summary. So the one action that would expose the fabrication is the action the feature exists to let you skip. The false statement is protected by the feature’s own value proposition.
The preview is the same failure in a different place. A preview is the checking step. It is where you go to confirm the thing is right before it goes out. When the preview renders the bug faithfully instead of catching it, the safety net has been checked and it reported clean.
The score is the sharpest version of all three. Its purpose is to spare you the résumé — that is what a screening score is for, and at scale it is a reasonable thing to want. To catch “14.4 years” you would have to open the CV and count the dates yourself, which is precisely the work the number exists to replace. And unlike the summary, the score arrives before you have decided to ask for anything: it is generated on upload, by a bot, and it is sitting on the candidate’s header the first time you ever see their name.
And in the summary’s case, the consequence is not cosmetic. “Applied” and “scorecards updated” are exactly the facts a hiring decision turns on. A busy manager reading that summary concludes this person applied for the role and has already been evaluated. Neither is true. That is a candidate advancing — or being cut — on a history that was generated rather than recorded, in a system whose whole job is to be the record.
Breezy’s own terms already say this can happen
I found this after the fact, and it changes how I read the whole thing.
Breezy’s terms of service carry a section headed AI DISCLOSURE. It says, in the company’s own words:
AI-generated content may be inaccurate, incomplete, misleading, outdated, biased, or otherwise unsuitable for its intended purpose. All content should be reviewed and independently verified before it is used, published, or distributed, including for factual accuracy […] AI-generated content should not be used as the sole basis for any recruiting or employment decision.
That document is dated 10 August 2026. I opened my account on the 11th. So on the day I started, the company had already written down, in a binding document, the exact failure mode this article spent three sections documenting.
I want to be careful about what that does and does not mean.
It does not mean the fabrication is acceptable, and I am not going to argue that a disclaimer settles the question. A summary that reports two events which never occurred is not “inaccurate” in the way a rounding error is inaccurate; it is a record of things that did not happen, produced by a system whose job is to hold the record.
But it does mean the behaviour is not a surprise to anyone inside the company. Nobody there has to be told. It is written down.
What is worth sitting with is where it is written down. The warning is in the terms of service — a document that is agreed to once, at signup, by clicking a box, and then never opened again. The features it warns about sit in the candidate header, in the summary panel, in the job description editor: three places you look at every working day. Not one of them repeats the warning.
And the instruction the terms give you — review and independently verify everything before you use it — is, for the summary, exactly the raw activity log I had to open to catch the fabrication. Which is the work the summary exists to save you. The terms tell you to do the thing the product is sold on not doing.
What to do about it
If you use an ATS with these features, three rules cost you very little:
- Do not treat an AI summary as a record. It is a suggestion about a record. Before it informs any decision about a person, open the raw activity log. Yes, that removes most of the time the summary saved you — that is the actual price of the feature.
- Read AI-written job posts in full before publishing. Read the body text, not the preview. The preview will not catch a placeholder; I have watched it not catch one.
- Check the AI toggles on every position, not once on the account. Applicant Insight and Resume Audit run on upload, without asking, and they keep running after a downgrade to the free plan. If Move Based on Score is also on, a generated number can move a real person through your pipeline while nobody is watching.
- Make all of those team rules, not personal habits. The person who publishes the job post is often not the person who generated it, and the person reading the summary is usually not the person who did the work being summarised. That gap is where this lands.
Is Breezy still worth using?
Yes, with the AI switched off. Underneath these features it is a competent, pleasant pipeline tool — the kanban board works the way everyone already expects, and the parts of the product that record what you did are accurate. It is specifically the layer that generates text on top of that record which I would not rely on.
If you are migrating from a spreadsheet, the import has its own set of problems worth reading before you commit. And if you are planning to sit on the free tier to avoid all of this, note that the AI does not stop when the paying does — I downgraded this same account and watched it score another candidate, unprompted, on a plan costing nothing.
Updated 22 August 2026 with the scoring results. The original piece covered the activity summary and the job description writer, both tested on 11 August 2026.
A second product read the same résumé and got the figure right — and then kept no copy of it. The comparison is here.
Disclosure: the Breezy HR link in this post pays me a commission if you sign up through it, and it is marked where it appears above. It costs you nothing extra, it did not change a word of what this post says, and I do not accept payment in exchange for favourable coverage. Every other link here is an ordinary one.