How to Write Fairer Performance Reviews With AI

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Write Fairer Performance Reviews With AI.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Write Fairer Performance Reviews With AI.

Use AI to check your reviews for bias, not to make the judgement. Gather evidence from the whole review period, rate each person yourself against written criteria, then ask an assistant to flag vague, personality-based or recent-only comments and compare wording across the team. Ratings, pay and promotion decisions stay with people.

The trap is that AI can make an unfair review more convincing. It turns a biased judgement into fluent, professional prose, and it only knows the evidence you hand it: if your notes cover the last six weeks, so will the review, however polished it reads. Fairness comes from the inputs and the checks. The drafting help is the smallest part of the gain.

Follow me on Instagram@sagnikteaches

Where unfairness creeps into small-team reviews

In a team of eight or ten, one manager often reviews everyone, from memory, in a single week. That setup produces predictable distortions. Knowing which one you're guarding against tells you what to ask the AI to look for.

Connect on LinkedInSagnik Bhattacharya
DistortionHow it shows up in the reviewWhat an AI pass can checkWhat it can't
RecencyThe last project dominates the whole yearDates of cited evidence across the periodFind evidence you never recorded
Halo or hornsOne strength or weakness colours every ratingRatings that all move together despite mixed evidenceKnow which trait is really driving your view
Leniency or central tendencyEveryone "meets expectations"Rating spread against evidence strengthDecide the right distribution for you
SimilarityPeople like the manager get richer, warmer write-upsLength, detail and tone across reviewsTell you why you warm to someone
VisibilityRemote or quiet staff get fewer concrete examplesCount of specific examples per personRetrieve work nobody saw
Personality language"Abrasive", "bubbly", "not a culture fit"Flag trait words and suggest behaviour-based rewritesJudge whether the underlying concern is valid
Double standards"Assertive" for one person, "aggressive" for anotherDifferent words for similar behaviour across reviewsKnow the full context of each incident

The right-hand column is the important one. Every limit there is solved by a human habit, mostly the evidence log below.

Subscribe on YouTube@codingliquids

Keep an evidence log for the whole period

A review can only be as fair as its evidence. Spend ten minutes a month per person noting what happened: date, what they did, the effect, and where the evidence lives (a client email, a QA report, a delivery record). Keep it in a shared document or your HR system, one row per note.

Here's an illustrative extract for one in-house translator at a nine-person translation agency:

DateWhat happenedEffectEvidence
JanuaryBuilt the glossary for a new legal client (1,400 terms)Reviser queries on that account fell in FebruaryGlossary file, query log
MarchMissed a 17:00 deadline by 3 hours on a medical leafletClient warned, no penaltyDelivery log
AprilFlagged a source-text error in a contract before deliveryClient thanked PM by emailEmail, 12 April
SeptemberReviewed two junior colleagues' work each weekTheir QA scores improvedQA reports

At review time, ask the assistant to organise the log against your criteria and, crucially, to report the gaps:

Below is a year of evidence notes for one team member (initials
only) and our review criteria. Group the notes under each criterion.
For each criterion, say how many notes support it and which months
they come from. List any criterion with fewer than two notes, and
any run of three or more months with no notes at all.
Do not rate the person. Do not add information that isn't in the notes.

An illustrative reply for the translator above began: "Accuracy: 5 notes (Jan, Feb, Apr, Aug, Oct). Deadline reliability: 2 notes (Mar, Nov), both about late delivery; no notes record on-time delivery. Collaboration: 4 notes (Sep to Dec only). No notes at all for May, June and July." That last line is the one to act on. Before writing anything, the manager checked the delivery log for those months and found 41 jobs delivered on time, which changed the picture of "deadline reliability" completely.

Write the criteria before anyone writes a review

Vague criteria invite vague reviews. Write down, for each role, what is being assessed and what each rating level looks like in behaviour. You can draft these with AI in half an hour, and the method is close to building interview scorecards, covered in building interview questions and scorecards with AI.

A filled-in example for one criterion, for the agency's translators:

Deadline reliabilityWhat it looks like
1: BelowSeveral late deliveries without warning; PMs chase for status
2: DevelopingMostly on time; late jobs flagged, but often on the day
3: MeetsOn time for about 95% of jobs; risks flagged at least a day ahead
4: ExceedsOn time almost without exception; proposes realistic dates at quoting stage

The other translator criteria were accuracy (QA errors per 1,000 words against the team target), terminology and glossary use, client feedback, and support for colleagues. Project managers had different ones: on-time delivery across their portfolio, margin per project, client retention and supplier management. Where you have a number, use it. Where you don't, the behavioural anchors do the same job. If you track skills separately, a staff training matrix gives you another source of dated evidence.

Draft the assessment yourself, then run the fairness pass

Write your own draft first, including the rating. It can be rough. Then replace names with initials or role titles and run a fairness check:

You are checking a draft performance review for fairness and clarity.
You'll get: the review criteria, the evidence summary, and my draft.
Flag, quoting the exact words:
1. Claims with no specific example or evidence behind them.
2. Words describing personality or attitude rather than behaviour.
3. Statements that only draw on the last two months.
4. Any rating my written evidence doesn't support, or supports a
   higher or lower rating than I gave.
5. Wording that carries assumptions about gender, age, family,
   health or background.
For each flag, suggest a rewrite as: behaviour, then effect, then
what to do next period. Don't change the ratings.

A shortened, illustrative output for a project manager's review:

1. "Not always on top of things": no example given. Evidence summary shows 2 late projects out of 38. Suggest citing those two and the on-time rate.
2. "A bit defensive when challenged": personality description. Rewrite around a specific meeting and what was said.
3. "Struggled with the software rollout": the only evidence is from October and November; the review period is 12 months.
4. Rating of 2 for supplier management: evidence includes onboarding 9 new freelance translators and cutting average supplier response time. The evidence appears to support a 3.
5. "Detail-oriented": may be read as a personality trait, but it's supported by QA data. Consider keeping it with the figure.

Not every flag needs acting on. The fifth is a false alarm: "detail-oriented" is backed by numbers, so the manager kept it and added the figure. The fourth is the most valuable, because it caught a rating the manager had lowered after one frustrating week in November. The manager, not the AI, then decided to raise it.

Before and after: one paragraph rewritten

This is the kind of paragraph that shows up in reviews written from memory:

R is a bit abrasive with PMs and not really a team player. Her work is generally good but she needs to be more positive.

After the fairness pass and a look at the evidence log, the manager rewrote it:

In three handovers (March, June and September), R replied to project managers by questioning the deadline without proposing an alternative, which led to two escalations to the operations lead. Her accuracy is strong: 0.8 QA errors per 1,000 words against a team target of 1.5. Next period: when a deadline looks unworkable, reply within the hour with a revised date or a reduced scope.

The concern survived. What disappeared was the personality verdict ("abrasive", "not a team player", "more positive"), which R couldn't act on and could reasonably dispute. What appeared was a dated pattern, a measured strength and a clear next step.

Compare reviews side by side before anything is final

Unfairness is often invisible in a single review and obvious across a set. Once all drafts are written, give the assistant every anonymised review together and ask for a comparison:

Here are 9 anonymised draft reviews from the same period, labelled
P1 to P9, with each person's role and whether they work remotely.
For each review, count: words, specific dated examples, personality
or attitude words. Then list any pattern by role, working pattern
(remote/office) or reviewer. Quote the phrases that differ in tone
for similar behaviour. Don't comment on whether ratings are right.

In the illustrative agency, the answer contained a table and one uncomfortable line: the two remote translators averaged 1.5 dated examples per review, against 3.8 for the four office-based translators, and their reviews were 40% shorter. Nothing in the ratings looked wrong in isolation. Across the set, the pattern was plain: remote work was simply less visible to the reviewer.

A quick rule of thumb for small teams: reopen any review with fewer than two dated examples per criterion, and any group whose average examples per review is less than half that of another group. With nine people, a pattern like this is a prompt for a second look, not proof of bias. The fix is always the same: go back to the evidence, not back to the wording.

Summarising peer feedback without flattening it

If you collect peer or client feedback, an assistant can condense it, but default summaries have two habits that hurt fairness. They average away the minority view, and they let the longest contributor dominate. Ask for a summary that preserves both:

Summarise the peer feedback below for one person (initials only).
For each theme, say how many of the 5 contributors raised it and
include one short direct quote. List separately any point raised by
only one contributor. If one contributor wrote more than half the
total words, say so. Don't soften criticism or praise.

In the illustrative agency, a project manager received five pieces of peer feedback. One colleague wrote 600 words of criticism about a single disputed project; the other four wrote around 80 words each, mostly positive. A plain "summarise this" request produced a summary that was two-thirds negative. The structured version showed the real picture: four of five contributors praised her planning, one raised a specific dispute that the owner then looked into separately. Same feedback, very different review.

A nine-person translation agency, start to finish

Here's how the illustrative agency ran its annual reviews with this method. The team: six in-house translators (two remote), two project managers and one vendor manager, reviewed by the owner and the operations lead.

  • Evidence logs: 10 minutes per person per month, about 18 hours across the year for both reviewers. This is extra time compared with the old approach, spread thinly.
  • Criteria: 90 minutes with an assistant to draft anchors for three roles, then an hour of editing.
  • Drafts: about 75 minutes per review, down from roughly 3 hours when the reviewers had to reconstruct the year from memory and inboxes.
  • Fairness pass: 15 minutes per review, including deciding which flags to accept. Across nine reviews it produced 47 flags; the reviewers acted on 29.
  • Calibration comparison: one hour for the set, which led to reopening three reviews.

Total review-week effort fell from about 27 hours to about 16, while the year-round logging added roughly 18. So the honest summary is "similar total time, much better evidence". Two ratings changed after the calibration step: one remote translator moved from "meets" to "exceeds" once her QA figures and on-time record were pulled in, and one project manager's supplier rating went up, as described above. One rating went down: an office-based translator whose warm write-up hadn't mentioned four late deliveries recorded in the log.

Lines the AI must not cross

Keep these decisions strictly human, and say so in your process notes:

  • Ratings, pay, bonuses, promotions, performance improvement plans and dismissals. The assistant can check whether your evidence supports a decision; it doesn't make it.
  • Anything about health, disability, pregnancy or family circumstances. Keep it out of prompts entirely. If absence is relevant, take advice on how it may fairly be considered.
  • Identifiable personal data in consumer tools. Use a business plan where content isn't used for training by default, and anonymise anyway. Deciding what may go where is covered in AI bias in small business decisions, alongside the wider fairness risks.

If you sell to customers in the EU or employ people there, note that the EU AI Act lists AI systems intended to monitor and evaluate the performance and behaviour of workers among its high-risk uses (Annex III). Those obligations for stand-alone systems were deferred to 2 December 2027. Using a chat assistant to tidy wording you've written is a different thing from buying a system that scores staff, but if you're considering a dedicated performance-scoring product, ask an employment or data-protection adviser before you buy.

Performance platforms are adding their own AI too. Lattice, for example, offers a Writing Assist feature for review responses and an AI feature that summarises the feedback written about someone in the current cycle. The same rules apply: summaries are only as balanced as the feedback collected, and the reviewer still owns the judgement.

How to tell the process is fairer

Pick a few measures and look at them each cycle:

  • Evidence depth: average dated examples per review, by group (remote and office, full-time and part-time, role).
  • Rating spread: whether ratings cluster at "meets" for everyone, which usually means the criteria aren't doing their job.
  • Staff view: one question in your next staff survey, "My review reflected my work across the whole year", on a five-point scale. Analysing those answers is covered in running staff surveys with AI analysis.
  • Disputes: the number of reviews challenged or reopened after the conversation.

With a small team, treat these as conversation starters. A single year's numbers from nine people won't prove anything statistically, but a remote-versus-office gap that repeats two cycles running tells you something real about how work is seen.

Preparing for the review conversation

The written review is half the job. AI is useful for rehearsal: paste your final review (anonymised) and ask, "What three questions is this person most likely to ask, and which parts of my review could they reasonably dispute?" For R's review above, the illustrative answer included "Which handovers? Were other translators given the same deadline?", which prompted the manager to bring the three dated examples and the team's on-time figures to the meeting.

Don't let the assistant script the conversation itself. People can tell when they're being read to, and the most useful part of a review is the part you didn't plan: what the person says about the year from their side. Write that down afterwards, and it becomes the first entry in next year's evidence log. If the review raises a policy question, such as how on-call time is recognised, check the answer against your employee handbook so what you say matches what's written.

Questions managers ask about AI and reviews

Can AI suggest the rating if I give it all the evidence?

It can, but don't ask it to. Once you've seen a suggested rating you tend to anchor on it, and the model has no way to weigh context it wasn't given, such as a difficult client or a mid-year change of role. Decide the rating yourself first, then ask the assistant whether your written evidence supports it. That order keeps the judgement yours.

Should I tell staff that AI helped with their review?

Yes, briefly. Say that you wrote the assessment and decided the rating, and that an AI tool helped check the wording for clarity and consistency. Hiding it risks trust if it comes out later, and being open signals that the tool was used to make reviews fairer rather than to save effort on them.

How often should the evidence log be updated?

Monthly works for most small teams: ten minutes per person, straight after a one-to-one if you hold them. Weekly is better for fast-moving roles, but a log that's kept monthly all year beats a weekly log abandoned in March. Set a recurring calendar reminder and treat an empty month as a prompt to ask the person what they worked on.

Further reads

Sources: EU AI Act Annex III point 4 (employment and workers' management); Lattice help pages on Writing Assist and feedback summaries; OpenAI and Anthropic business plan data-use terms.

Want your review process checked for fairness?

On a 1:1 call we'll look at how your managers gather evidence, set up the fairness and calibration prompts, and agree which parts of the review stay strictly human.

Book a 1:1 call with me