When AI Makes Work Visible, Evaluation Design Decides Fairness
AI adoption always brings a kind of visibility — case volume, the details of how work gets handled. Whether that change becomes a threat or a path to fairness comes down to how the evaluation is designed, in our experience.
When we talk with clients about deploying AI, one concern comes up almost every time: bring AI into day-to-day operations, and everyone’s work becomes visible — won’t that create pushback?
Case volume, time spent per task, the details of an exchange — once AI sits inside a workflow, information that used to live only in someone’s judgment or private notes starts showing up as a log. This article lays out how we think about that side effect of visibility: not only as a threat, but as something worth designing for.
Why visibility becomes a common objection
When we explain an AI rollout, it’s not unusual to hear reactions like “now everyone will see if I’m slacking off” or “this will end up being used to judge me.” That reaction is not off base. Once AI enters daily work, case volume, response time, and the details of how something was handled do start showing up in far finer-grained records than before.
The effect is strongest in work that had been highly personal, where methods and speed varied a lot from person to person — this is often the first time a genuine basis for comparison exists. Differences that used to be waved away with “well, they’re the veteran” now sit side by side in the record. At the core of the unease, as we see it, is this new possibility of being compared.
When that worry turns out to be justified
In our experience, though, whether the worry becomes reality has less to do with deploying AI itself than with how the evaluation is designed afterward. Visibility tends to turn straight into a threat under a particular kind of design: one that looks only at a simple number, like case count, and ignores how difficult each case actually was. Under that design, the people handling the most exceptions and the trickiest requests come out looking worse simply because their numbers are lower.
The same problem shows up when a single, uniform metric is applied without regard to department or individual circumstances. Measure a new hire and a veteran, or someone doing mostly routine work and someone handling mostly exceptions, on the same yardstick, and distortion is inevitable somewhere. When records like this are made visible on top of a design this rough, you get the inversion where the people working hardest end up looking the worst on paper — and the resulting distrust on the ground is entirely justified. This is exactly why we don’t think unease about visibility should be dismissed as mere sentiment.
When visibility becomes a path to fairness
At the same time, we’ve also seen the same visibility work in the opposite direction once the evaluation design is done carefully. Plenty of cases exist where careful handling and quiet, steady adjustments went unrecorded and effectively invisible as effort. Once records are read alongside the context and difficulty of each case, the work of people who were doing a genuinely good job starts to come through more fairly than before — evaluation stops being a matter of reputation and impression and becomes something backed by the actual exchanges themselves.
One thing we’ve come to believe from working alongside clients is that AI adoption surfaces problems that were already hiding inside an organization more than it improves efficiency on its own. Who was actually carrying which workload, and how much the evaluation standard varied by department or manager — these were facts about the organization that existed long before AI arrived. Visibility doesn’t create that reality; it simply brings into view what was already there but unseen. Seen that way, visibility can become a welcome byproduct for an organization rather than a threat to it.
One more thing we keep in mind on the ground: resistance to visibility shouldn’t be written off as simple irrational aversion to change. Some of that resistance reflects legitimate circumstances the evaluator never had full sight of, not just inertia. A case in point: someone who looks slower on response-time metrics may in fact be spending more time explaining things carefully, and get higher marks from customers as a result. Differences like that get lost in an evaluation design that reads records only from a single angle. Whether an organization can weigh the pushback against actual outcomes, rather than dismiss it, is part of what determines whether visibility ends up serving fairness.
Our conclusion: the danger isn’t visibility itself
Taken together, our conclusion comes down to one point. What actually deserves caution in an AI rollout isn’t visibility as a phenomenon — it’s exposing the record to the workplace before the evaluation design behind it has been worked out. We put weight on treating visibility not as something that simply happens and has to be managed after the fact, but as something a company can design for in advance.
Trying to stop records from being kept, or keeping them deliberately vague, doesn’t solve anything at the root. It only postpones the real issue that visibility brings to the surface: the unevenness in how work had been done and in evaluation standards that already existed.
A practical starting point: what to settle before turning on visibility
So this isn’t just something to read and set aside, here is one starting point for thinking through an evaluation design. This is only an example, not a prescription for the “correct” design — the content and priority of each item should be filled in by each organization according to the nature of its work, its size, and its existing evaluation culture.
| Consideration | Question to settle |
|---|---|
| What gets measured | Beyond raw count and time, how is case difficulty and context reflected? |
| How to adjust for it | Is there a mechanism so people handling more exceptions or complex requests aren’t penalized? |
| Who reviews it | Is there a person who actually looks at the content, rather than just reading off the numbers? |
| How feedback reaches people | Are both good marks and areas to improve communicated back to the person? |
| Review cadence | Who revisits the evaluation standard itself, and how often? |
If most of these boxes stay blank when you try to fill them in, and all you have is “well, at least the record exists” — in our view, that’s a sign visibility has started without any design behind it.
Having someone assigned to review the records is not, by itself, enough to feel reassured either. If that reviewer just rubber-stamps the numbers without actually looking at the content, the evaluation becomes a formality and the benefit of visibility disappears. What matters isn’t whether a person is in the loop, but what that person is actually checking.
Addressing the pushback: “isn’t this just more surveillance?”
A reasonable objection to everything above is that visibility is still, in the end, just more surveillance. We think that objection has a point. The mechanism of keeping records can become either surveillance or evaluation, depending entirely on how it’s used. What decides which one it becomes is the operator’s stance: is the record being used to constrain behavior, or to understand what actually happened and put that to use going forward? Lean toward the former, and people on the ground become guarded, and keeping the record starts to feel like an end in itself. Lean toward the latter, and visibility becomes material that rewards people who were doing solid work, and a way to surface organizational problems that had never been sorted out.
Choosing to soften the records, or simply not use them for evaluation, out of fear of visibility, might look like the safe choice — but it does nothing more than preserve the existing unevenness of a personalized evaluation system. In our view, there is more value in carefully working out the evaluation design that follows visibility than in avoiding visibility itself. When considering an AI rollout, we’d recommend designing not only for what becomes visible, but for how what becomes visible will be handled.