Use AI to check your reviews for bias, not to make the judgement. Gather evidence from the whole review period, rate each person yourself against written criteria, then ask an assistant to flag vague, personality-based or recent-only comments and compare wording across the team. Ratings, pay and promotion decisions stay with people.
The trap is that AI can make an unfair review more convincing. It turns a biased judgement into fluent, professional prose, and it only knows the evidence you hand it: if your notes cover the last six weeks, so will the review, however polished it reads. Fairness comes from the inputs and the checks. The drafting help is the smallest part of the gain.
Where unfairness creeps into small-team reviews
In a team of eight or ten, one manager often reviews everyone, from memory, in a single week. That setup produces predictable distortions. Knowing which one you're guarding against tells you what to ask the AI to look for.
| Distortion | How it shows up in the review | What an AI pass can check | What it can't |
|---|---|---|---|
| Recency | The last project dominates the whole year | Dates of cited evidence across the period | Find evidence you never recorded |
| Halo or horns | One strength or weakness colours every rating | Ratings that all move together despite mixed evidence | Know which trait is really driving your view |
| Leniency or central tendency | Everyone "meets expectations" | Rating spread against evidence strength | Decide the right distribution for you |
| Similarity | People like the manager get richer, warmer write-ups | Length, detail and tone across reviews | Tell you why you warm to someone |
| Visibility | Remote or quiet staff get fewer concrete examples | Count of specific examples per person | Retrieve work nobody saw |
| Personality language | "Abrasive", "bubbly", "not a culture fit" | Flag trait words and suggest behaviour-based rewrites | Judge whether the underlying concern is valid |
| Double standards | "Assertive" for one person, "aggressive" for another | Different words for similar behaviour across reviews | Know the full context of each incident |
The right-hand column is the important one. Every limit there is solved by a human habit, mostly the evidence log below.
Keep an evidence log for the whole period
A review can only be as fair as its evidence. Spend ten minutes a month per person noting what happened: date, what they did, the effect, and where the evidence lives (a client email, a QA report, a delivery record). Keep it in a shared document or your HR system, one row per note.
Here's an illustrative extract for one in-house translator at a nine-person translation agency:
| Date | What happened | Effect | Evidence |
|---|---|---|---|
| January | Built the glossary for a new legal client (1,400 terms) | Reviser queries on that account fell in February | Glossary file, query log |
| March | Missed a 17:00 deadline by 3 hours on a medical leaflet | Client warned, no penalty | Delivery log |
| April | Flagged a source-text error in a contract before delivery | Client thanked PM by email | Email, 12 April |
| September | Reviewed two junior colleagues' work each week | Their QA scores improved | QA reports |
At review time, ask the assistant to organise the log against your criteria and, crucially, to report the gaps:
Below is a year of evidence notes for one team member (initials
only) and our review criteria. Group the notes under each criterion.
For each criterion, say how many notes support it and which months
they come from. List any criterion with fewer than two notes, and
any run of three or more months with no notes at all.
Do not rate the person. Do not add information that isn't in the notes.
An illustrative reply for the translator above began: "Accuracy: 5 notes (Jan, Feb, Apr, Aug, Oct). Deadline reliability: 2 notes (Mar, Nov), both about late delivery; no notes record on-time delivery. Collaboration: 4 notes (Sep to Dec only). No notes at all for May, June and July." That last line is the one to act on. Before writing anything, the manager checked the delivery log for those months and found 41 jobs delivered on time, which changed the picture of "deadline reliability" completely.
Write the criteria before anyone writes a review
Vague criteria invite vague reviews. Write down, for each role, what is being assessed and what each rating level looks like in behaviour. You can draft these with AI in half an hour, and the method is close to building interview scorecards, covered in building interview questions and scorecards with AI.
A filled-in example for one criterion, for the agency's translators:
| Deadline reliability | What it looks like |
|---|---|
| 1: Below | Several late deliveries without warning; PMs chase for status |
| 2: Developing | Mostly on time; late jobs flagged, but often on the day |
| 3: Meets | On time for about 95% of jobs; risks flagged at least a day ahead |
| 4: Exceeds | On time almost without exception; proposes realistic dates at quoting stage |
The other translator criteria were accuracy (QA errors per 1,000 words against the team target), terminology and glossary use, client feedback, and support for colleagues. Project managers had different ones: on-time delivery across their portfolio, margin per project, client retention and supplier management. Where you have a number, use it. Where you don't, the behavioural anchors do the same job. If you track skills separately, a staff training matrix gives you another source of dated evidence.
Draft the assessment yourself, then run the fairness pass
Write your own draft first, including the rating. It can be rough. Then replace names with initials or role titles and run a fairness check:
You are checking a draft performance review for fairness and clarity.
You'll get: the review criteria, the evidence summary, and my draft.
Flag, quoting the exact words:
1. Claims with no specific example or evidence behind them.
2. Words describing personality or attitude rather than behaviour.
3. Statements that only draw on the last two months.
4. Any rating my written evidence doesn't support, or supports a
higher or lower rating than I gave.
5. Wording that carries assumptions about gender, age, family,
health or background.
For each flag, suggest a rewrite as: behaviour, then effect, then
what to do next period. Don't change the ratings.
A shortened, illustrative output for a project manager's review:
1. "Not always on top of things": no example given. Evidence summary shows 2 late projects out of 38. Suggest citing those two and the on-time rate.
2. "A bit defensive when challenged": personality description. Rewrite around a specific meeting and what was said.
3. "Struggled with the software rollout": the only evidence is from October and November; the review period is 12 months.
4. Rating of 2 for supplier management: evidence includes onboarding 9 new freelance translators and cutting average supplier response time. The evidence appears to support a 3.
5. "Detail-oriented": may be read as a personality trait, but it's supported by QA data. Consider keeping it with the figure.
Not every flag needs acting on. The fifth is a false alarm: "detail-oriented" is backed by numbers, so the manager kept it and added the figure. The fourth is the most valuable, because it caught a rating the manager had lowered after one frustrating week in November. The manager, not the AI, then decided to raise it.
Before and after: one paragraph rewritten
This is the kind of paragraph that shows up in reviews written from memory:
R is a bit abrasive with PMs and not really a team player. Her work is generally good but she needs to be more positive.
After the fairness pass and a look at the evidence log, the manager rewrote it:
In three handovers (March, June and September), R replied to project managers by questioning the deadline without proposing an alternative, which led to two escalations to the operations lead. Her accuracy is strong: 0.8 QA errors per 1,000 words against a team target of 1.5. Next period: when a deadline looks unworkable, reply within the hour with a revised date or a reduced scope.
The concern survived. What disappeared was the personality verdict ("abrasive", "not a team player", "more positive"), which R couldn't act on and could reasonably dispute. What appeared was a dated pattern, a measured strength and a clear next step.
Compare reviews side by side before anything is final
Unfairness is often invisible in a single review and obvious across a set. Once all drafts are written, give the assistant every anonymised review together and ask for a comparison:
Here are 9 anonymised draft reviews from the same period, labelled
P1 to P9, with each person's role and whether they work remotely.
For each review, count: words, specific dated examples, personality
or attitude words. Then list any pattern by role, working pattern
(remote/office) or reviewer. Quote the phrases that differ in tone
for similar behaviour. Don't comment on whether ratings are right.
In the illustrative agency, the answer contained a table and one uncomfortable line: the two remote translators averaged 1.5 dated examples per review, against 3.8 for the four office-based translators, and their reviews were 40% shorter. Nothing in the ratings looked wrong in isolation. Across the set, the pattern was plain: remote work was simply less visible to the reviewer.
A quick rule of thumb for small teams: reopen any review with fewer than two dated examples per criterion, and any group whose average examples per review is less than half that of another group. With nine people, a pattern like this is a prompt for a second look, not proof of bias. The fix is always the same: go back to the evidence, not back to the wording.
Summarising peer feedback without flattening it
If you collect peer or client feedback, an assistant can condense it, but default summaries have two habits that hurt fairness. They average away the minority view, and they let the longest contributor dominate. Ask for a summary that preserves both:
Summarise the peer feedback below for one person (initials only).
For each theme, say how many of the 5 contributors raised it and
include one short direct quote. List separately any point raised by
only one contributor. If one contributor wrote more than half the
total words, say so. Don't soften criticism or praise.
In the illustrative agency, a project manager received five pieces of peer feedback. One colleague wrote 600 words of criticism about a single disputed project; the other four wrote around 80 words each, mostly positive. A plain "summarise this" request produced a summary that was two-thirds negative. The structured version showed the real picture: four of five contributors praised her planning, one raised a specific dispute that the owner then looked into separately. Same feedback, very different review.
A nine-person translation agency, start to finish
Here's how the illustrative agency ran its annual reviews with this method. The team: six in-house translators (two remote), two project managers and one vendor manager, reviewed by the owner and the operations lead.
- Evidence logs: 10 minutes per person per month, about 18 hours across the year for both reviewers. This is extra time compared with the old approach, spread thinly.
- Criteria: 90 minutes with an assistant to draft anchors for three roles, then an hour of editing.
- Drafts: about 75 minutes per review, down from roughly 3 hours when the reviewers had to reconstruct the year from memory and inboxes.
- Fairness pass: 15 minutes per review, including deciding which flags to accept. Across nine reviews it produced 47 flags; the reviewers acted on 29.
- Calibration comparison: one hour for the set, which led to reopening three reviews.
Total review-week effort fell from about 27 hours to about 16, while the year-round logging added roughly 18. So the honest summary is "similar total time, much better evidence". Two ratings changed after the calibration step: one remote translator moved from "meets" to "exceeds" once her QA figures and on-time record were pulled in, and one project manager's supplier rating went up, as described above. One rating went down: an office-based translator whose warm write-up hadn't mentioned four late deliveries recorded in the log.
Lines the AI must not cross
Keep these decisions strictly human, and say so in your process notes:
- Ratings, pay, bonuses, promotions, performance improvement plans and dismissals. The assistant can check whether your evidence supports a decision; it doesn't make it.
- Anything about health, disability, pregnancy or family circumstances. Keep it out of prompts entirely. If absence is relevant, take advice on how it may fairly be considered.
- Identifiable personal data in consumer tools. Use a business plan where content isn't used for training by default, and anonymise anyway. Deciding what may go where is covered in AI bias in small business decisions, alongside the wider fairness risks.
If you sell to customers in the EU or employ people there, note that the EU AI Act lists AI systems intended to monitor and evaluate the performance and behaviour of workers among its high-risk uses (Annex III). Those obligations for stand-alone systems were deferred to 2 December 2027. Using a chat assistant to tidy wording you've written is a different thing from buying a system that scores staff, but if you're considering a dedicated performance-scoring product, ask an employment or data-protection adviser before you buy.
Performance platforms are adding their own AI too. Lattice, for example, offers a Writing Assist feature for review responses and an AI feature that summarises the feedback written about someone in the current cycle. The same rules apply: summaries are only as balanced as the feedback collected, and the reviewer still owns the judgement.
How to tell the process is fairer
Pick a few measures and look at them each cycle:
- Evidence depth: average dated examples per review, by group (remote and office, full-time and part-time, role).
- Rating spread: whether ratings cluster at "meets" for everyone, which usually means the criteria aren't doing their job.
- Staff view: one question in your next staff survey, "My review reflected my work across the whole year", on a five-point scale. Analysing those answers is covered in running staff surveys with AI analysis.
- Disputes: the number of reviews challenged or reopened after the conversation.
With a small team, treat these as conversation starters. A single year's numbers from nine people won't prove anything statistically, but a remote-versus-office gap that repeats two cycles running tells you something real about how work is seen.
Preparing for the review conversation
The written review is half the job. AI is useful for rehearsal: paste your final review (anonymised) and ask, "What three questions is this person most likely to ask, and which parts of my review could they reasonably dispute?" For R's review above, the illustrative answer included "Which handovers? Were other translators given the same deadline?", which prompted the manager to bring the three dated examples and the team's on-time figures to the meeting.
Don't let the assistant script the conversation itself. People can tell when they're being read to, and the most useful part of a review is the part you didn't plan: what the person says about the year from their side. Write that down afterwards, and it becomes the first entry in next year's evidence log. If the review raises a policy question, such as how on-call time is recognised, check the answer against your employee handbook so what you say matches what's written.
Questions managers ask about AI and reviews
Can AI suggest the rating if I give it all the evidence?
It can, but don't ask it to. Once you've seen a suggested rating you tend to anchor on it, and the model has no way to weigh context it wasn't given, such as a difficult client or a mid-year change of role. Decide the rating yourself first, then ask the assistant whether your written evidence supports it. That order keeps the judgement yours.
Should I tell staff that AI helped with their review?
Yes, briefly. Say that you wrote the assessment and decided the rating, and that an AI tool helped check the wording for clarity and consistency. Hiding it risks trust if it comes out later, and being open signals that the tool was used to make reviews fairer rather than to save effort on them.
How often should the evidence log be updated?
Monthly works for most small teams: ten minutes per person, straight after a one-to-one if you hold them. Weekly is better for fast-moving roles, but a log that's kept monthly all year beats a weekly log abandoned in March. Set a recurring calendar reminder and treat an empty month as a prompt to ask the person what they worked on.
Further reads
- How to Spot Bias and Stereotypes in AI-Written Content — More wording checks that catch stereotypes in any AI-assisted text.
- AI for HR in Small Businesses: 12 Tasks You Can Hand Over Safely — Other HR jobs that are safe to hand to AI.
- How to Classify Business Data Before Using AI Tools — Decide which staff data may go into which tool.
- Can AI Screen CVs Fairly? What Small Employers Need to Know — The same fairness questions, applied at the hiring stage.
- How to Handle Staff Who Over-Rely on AI — What to do when managers let AI write reviews unchecked.
- How to Set Up Human Review for AI Work Without Slowing Down — A general review routine for AI-assisted documents.
- AI Employee Onboarding: Plan a New Starter's First Two Weeks — Use AI to draft a realistic day-by-day plan for a new starter's first fortnight, a reading path through your documents, and an assistant they can ask.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: EU AI Act Annex III point 4 (employment and workers' management); Lattice help pages on Writing Assist and feedback summaries; OpenAI and Anthropic business plan data-use terms.