ChatGPT for Performance Reviews
Most advice worries ChatGPT reviews will sound robotic. The bigger risk is invented specifics smoothed into confident-sounding fiction. These chatgpt performance review prompts fix that instead.
Write a performance review for [employee name], who is a solid performer on my team.
Most advice about ChatGPT and performance reviews is wrong about the risk. Everyone worries about the review sounding robotic or generic — a fair concern, but a minor one. The bigger risk almost nobody talks about is ChatGPT confidently reframing a specific, factual example you gave it into something vaguer and less accurate, because the model naturally smooths language toward safer, more generic phrasing unless you explicitly stop it from doing that.
Before: The Weak Prompt
Write a performance review for [employee name], who is a solid performer on my team.This asks ChatGPT to generate a review almost entirely from its own imagination, since it has no actual information about what the employee did this quarter. The output will be generically positive, filled with phrases like "consistently delivers quality work" and "shows strong initiative" — plausible-sounding, completely unverifiable, and disconnected from anything the employee actually did.
This kind of prompt is tempting precisely because it's fast, and a busy manager staring down eight reviews due by Friday can talk themselves into believing a quick generic pass is good enough, especially for an employee they genuinely do think is doing well. But "doing well" without specifics is exactly the feedback that leaves an employee unsure what to keep doing, since there's nothing concrete in the review to point back to and repeat. It also leaves the manager with nothing solid to reference months later, whether for a promotion conversation, a compensation discussion, or simply their own memory of what actually happened that quarter.
Why It Fails
A performance review that could apply to almost any "solid performer" on any team isn't a performance review — it's a compliment template. Employees can tell the difference between specific, earned feedback and generic praise, and generic reviews tend to erode trust in the review process over time rather than building it, because they signal that the manager either doesn't know or didn't take the time to note what the employee actually did.
There's also a compounding effect across review cycles that's easy to miss in the moment. An employee who receives one generic review might not think much of it. An employee who receives the same generic phrases quarter after quarter starts to reasonably conclude that the review process itself is a formality nobody takes seriously, which undermines the credibility of the one time a review actually does need to carry real, specific weight — a promotion case, a performance improvement conversation, or a compensation decision that depends on evidence a string of generic past reviews simply doesn't provide.
After: The Improved Prompt
Here are specific notes on [employee name]'s work this quarter: [paste specific examples, projects, outcomes]
Draft a performance review that:
- References only the specific examples I've provided, not generic strengths
- For each strength, ties it to a specific project or outcome from my notes
- For the growth area, uses a specific example, not a vague generalization
Do not add any accomplishments, projects, or qualities I haven't mentioned in my notes.What this does: The explicit "do not add" instruction is the critical safeguard here — without it, ChatGPT will often fill perceived gaps in your notes with plausible-sounding but invented specifics, which is a genuinely dangerous failure mode in a document that could affect someone's compensation or career trajectory.
⚠️ Common mistake: Giving ChatGPT vague notes ("did well on the Johnson project") and trusting it to fill in specifics. The model will generate something that sounds specific — a particular skill demonstrated, a particular outcome achieved — but it's inventing those details based on what's plausible for a project with that name, not reporting anything you actually observed.
Breaking Down Each Element
Providing your own notes rather than a general characterization ("solid performer") shifts the entire exercise from generation to organization — you're asking ChatGPT to structure and phrase information you already have, not to generate the substance of the review itself. The instruction to tie each strength to a specific example forces concreteness in the final language, preventing drift back toward generic phrasing even when your original notes were specific, since language models have a mild but persistent pull toward smoothing specific details into more generic, safer-sounding statements during rewriting. The growth area instruction matters especially — vague growth feedback ("could improve communication") is nearly useless to an employee trying to actually change something, while a specific example gives them something concrete to work on.
This smoothing tendency is worth understanding on its own, since it shows up well beyond performance reviews. Language models tend to favor phrasing that reads as broadly safe and professionally acceptable, which is generally a reasonable default but works directly against the goal of a document whose entire value comes from being specific and concrete. Re-reading a ChatGPT-drafted review against your original raw notes, checking that nothing specific got quietly rounded off into something vaguer, is a habit worth building regardless of how detailed your original prompt was.
Real-World Scenario: A Manager Writing Reviews for a Team of Eight
Devon manages a team of eight and needed to write eight performance reviews in a single week — a volume that made shortcuts genuinely tempting, and exactly the situation where generic, templated language creeps in fastest.
Here are my raw notes on [employee]: [paste bullet-point notes, however rough]
Turn these into a coherent performance review draft, but keep every specific detail from my notes — don't smooth them into general statements.
If any of my notes are too vague to turn into a specific point, list them separately as things I should add more detail to before finalizing.What this does: The "list separately as things needing more detail" instruction turns a potential blind spot into a visible checklist, which is far more useful under time pressure than ChatGPT silently smoothing over a vague note into a confident-sounding but ultimately unverified statement.
⚡ Pro tip: Write your raw notes as bullet points immediately after the events happen throughout the quarter, not from memory right before review season. Reviews built from notes taken in the moment are consistently more specific and accurate than reviews built from a manager trying to reconstruct a whole quarter's worth of detail in one sitting under deadline pressure.
Real-World Scenario: An HR Business Partner Standardizing Review Language
Tasha works in HR and noticed inconsistent review quality across managers — some wrote detailed, specific reviews, others wrote generic ones, and the inconsistency was creating fairness concerns across the organization when reviews were compared during calibration discussions.
Here is a manager's rough notes on an employee: [paste notes]
Here is our company's rubric for performance ratings: [paste rubric]
Draft a review that maps the manager's specific notes to the relevant rubric categories.
Flag any rubric category where the manager's notes don't provide enough specific detail to support a rating, so the manager knows what to add before finalizing.What this does: Mapping specific manager notes against a formal rubric, and flagging gaps rather than filling them, helps standardize review quality across an organization without generating language that oversteps what any individual manager actually observed — a critical distinction in calibration processes where inconsistent evidence behind similar ratings can create real fairness and legal exposure concerns.
⚠️ Common mistake: Using ChatGPT to standardize review language across an entire organization without a human check on whether the underlying manager notes actually support the final rating. Consistent phrasing doesn't fix inconsistent underlying evidence — it just makes the inconsistency harder to spot on the surface.
Variations for Different Contexts
For self-reviews rather than manager reviews, the same "specific notes in, specific language out" principle applies, but it's worth explicitly asking the employee to quantify accomplishments wherever possible before drafting — self-reviews suffer from vagueness just as often as manager reviews do, often because employees underestimate their own accomplishments and default to modest, generic language that undersells specific, measurable work.
For reviews that need to address a genuine performance concern rather than general feedback, the specificity principle matters even more, since vague language in this context can create real ambiguity about whether a documented issue was ever actually communicated clearly. In these cases, it's worth explicitly asking ChatGPT to flag any sentence that could be read as ambiguous about what specifically needs to change and by when — a concern phrased as "needs to improve time management" leaves far more room for later disagreement about what was actually expected than one phrased around a specific, dated example and a specific, measurable expectation going forward.
Save and Reuse This
Keep a standard review-drafting prompt template — the "specific notes in, no invented details, flag the gaps" structure — saved and versioned in PromptABCD, since this discipline matters just as much in your tenth review cycle as your first, and it's exactly the kind of safeguard that's easy to skip when you're rushing through eight reviews in a week under deadline pressure.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
