Your Compliance Manager has had a quiet quarter. No licenses expired. No regulator called. No major contractual risk slipped through. No crisis required an emergency meeting.
So what exactly did they achieve?
Now take your Executive Assistant. Your calendar mostly worked. Important meetings were prepped. Conflicting commitments were resolved before they became your problem.
How do you measure that?
Or your HR Manager, who handled a difficult employee issue before it turned into a grievance, and identified a hiring problem before the wrong candidate joined. What number captures that?
Growing businesses eventually run into a performance-management problem that spreadsheets cannot solve. Some roles create obvious outputs—Sales generates revenue, Production creates units. But other roles create value through judgment, prevention, coordination, and reliability.
Their best work may not produce a large, visible event. Sometimes, their best work is precisely why the event never happened.
And that creates a dangerous temptation: If the work is difficult to measure, invent something countable. That is exactly where performance measurement starts destroying business value.
The Trap of "Precision Theatre"
When founders struggle to measure a role, they usually confuse measurement with counting.
To measure an Executive Assistant, they count: emails answered, meetings scheduled, or tasks closed. Now you have data. But do you actually know if you have a good Executive Assistant?
An EA could schedule 180 meetings in one month while allowing your time to be consumed by low-value conversations. They could answer every message quickly while missing the three that needed urgent attention. The numbers tell you that work happened. They do not tell you if value was created.
Numbers feel safe because they look precise. A manager rates an employee:
- Communication: 4/5
- Initiative: 3/5
- Strategic Thinking: 3.5/5
The spreadsheet calculates an overall score of 3.63. It looks scientific. But ask the manager: Why was initiative a 3 rather than a 4? Which specific situations are you basing that rating on? Often, there is no clear answer.
We have not eliminated subjectivity. We have just given subjectivity decimal places. That is Precision Theatre.
Instead of asking, "Can I put a number on this?" you must ask, "Does this evidence actually tell me if the role is working?"
The Firefighter Bias
There is a bias inside many businesses that rarely gets named: The person who solves a visible crisis gets remembered. The person who prevented the crisis gets ignored.
Imagine two Operations Managers.
- Manager A spends Thursday afternoon rescuing a major customer delivery. The wrong stock was prepped. The customer is furious. They arrange emergency transport and personally ensure delivery at 8 p.m. The founder applauds: "Excellent recovery. That's true ownership."
- Manager B noticed on Tuesday that the stock allocation was wrong. They corrected it, confirmed transport on Wednesday, and Thursday passed quietly. No crisis. Nobody applauds.
This is the Firefighter Bias. Visible rescue always looks more impressive than invisible prevention.
If your performance system only rewards visible output and dramatic intervention, you will accidentally teach your people to value activity over control, and rescue over prevention.
The Four Shapes of Value
Different roles create value differently. You cannot ask every role for the same shape of proof. Ask instead: Where does the value of this role become visible?
There are four common patterns:
1. Produce
Roles that create a clear work product (Analysts, Designers, Engineers).
- The Mistake: Measuring only volume (number of reports).
- The Evidence: Is the output accurate, timely, complete, and good enough to support the next business decision?
2. Enable
Roles that make somebody else more effective (Executive Assistants, IT Support, HR).
- The Mistake: Counting tasks (number of diary changes).
- The Evidence: Does important work move more smoothly because this role exists? Are conflicts anticipated? Are low-value tasks filtered out before hitting the founder's desk?
3. Protect
Roles that prevent deterioration (Compliance, Risk, Financial Control).
- The Mistake: Counting incidents (or lack thereof), assuming a quiet month was skill rather than luck.
- The Evidence: Were required controls actually performed? Were exceptions identified early? Did known risks sit unresolved until somebody else noticed them?
4. Decide
Roles that create value through judgment (Managers, Senior Specialists, Strategy).
- The Mistake: Trying to count "decisions made."
- The Evidence: How did they think? Did they gather relevant facts, understand trade-offs, consult the right people, and make a recommendation rather than just passing the problem upward?
Qualitative Does Not Mean Vague
Founders often think that if a metric isn't a hard number, it must be vague. That is false. Structured judgment is highly specific.
Suppose you tell a Finance Manager: "I expect good judgment." That is vague.
But if you establish a structured standard for an unusual payment request, you have an objective framework:
- Confirm the facts and identify the control exception.
- Establish the consequence of delaying vs. proceeding.
- Consult affected departments.
- Provide a clear recommendation to the founder.
When a crisis hits, you don't evaluate them with a lazy "3/5 for problem-solving." You evaluate them against the standard. Did they present options? Did they check the controls? That is structured judgment. It is entirely measurable, even without a percentage.
The Two Windows of Evidence
Hard-to-quantify roles need evidence from two different windows so that one heroic rescue (or one bad mistake) doesn't distort the entire year.
1. The Ordinary Work (System Reliability) Does work normally arrive when promised? Are handovers reliable? Do stakeholders continually need to chase them? A brilliant response to a crisis does not compensate for causing routine friction every single week.
2. The Revealing Moments (System Stress) What happens when a supplier fails? When a deadline moves suddenly? When a senior stakeholder makes an unusual request? Routine evidence tells you if the system works. Critical moments tell you if the person can think when the system is tested.
The "Evidence Test" (5 Questions)
Before inventing a fake KPI for a hard-to-measure role, run it through these five questions:
- What should be different in the business because this role exists? (If you can't answer this, you have a role-design problem, not a measurement problem).
- What shape does the value take? (Produce, Enable, Protect, or Decide?)
- What evidence would naturally exist if this outcome were happening? (Use real work output—reports, decision logs, lack of escalations—not manufactured HR forms).
- What situations reveal whether this person can actually perform?
- Could another reasonable manager look at this same evidence and understand why we rated them this way?
Stop Forcing Everything Into a Spreadsheet
There is comfort in seeing a dashboard full of percentages. It feels controlled. But performance management is not improved by turning everything human beings do into a number.
Trying to make qualitative roles look more objective than they really are makes your performance system less accurate. The alternative to numbers is not vague management. It is disciplined evidence.
The question should never be: "How do I force this role into a KPI?" The question must be: "What would we need to observe to know this role is creating the value we bought?"
Sometimes the strongest evidence of performance is that a problem which used to consume the founder's time quietly stopped happening. You do not need to invent a number to make that real.
