flowchart LR
A["Annual appraisal:<br>one cycle, one rater,<br>one rating"] --> B["Agile performance management"]
B --> C["Continuous conversations<br>(cadence)"]
B --> D["Peer and 360 input<br>(evidence base)"]
B --> E["Unbundled decisions<br>(development, evaluation, reward)"]
style A fill:#ffebee,stroke:#C62828
style C fill:#e3f2fd,stroke:#1976D2
style D fill:#e8f5e9,stroke:#388E3C
style E fill:#fff8e1,stroke:#F9A825
9 Agile Performance Management
You will be able to:
- Explain why the annual appraisal declined and what evidence drove leading firms to replace it.
- Design a system of real-time feedback and continuous performance conversations, including cadence, content, and manager capability.
- Deploy peer reviews and 360-degree feedback in ways that add perspective without adding politics.
- Separate the development, evaluation, and reward decisions that the traditional appraisal bundled together, and defend the design trade-offs this separation creates.
9.1 Introduction
No HR process illustrates the shift from traditional to agile more completely than performance management. Chapter 4 introduced the moment the traditional apparatus cracked: from around 2012, firms including Adobe, Deloitte, Microsoft, and eventually GE, the very inventor of forced ranking, dismantled annual ratings in favour of frequent, forward-looking conversations (Peter Cappelli & Anna Tavis, 2016). The reasons were practical before they were philosophical. Work had become project-based and fast-moving, so a yearly cycle described work that had finished months earlier; ratings consumed enormous management time, Deloitte counted roughly two million hours a year, while demonstrably failing to improve performance; and the backward-looking, judgment-heavy format was better at generating anxiety and appeal than at changing behaviour (Marcus Buckingham & Ashley Goodall, 2015).
Agile performance management is not the absence of management, the misreading Chapter 4 warned against, but a redesign around the principles of Part I: short feedback loops instead of annual batches, transparency of goals and progress, and the growth mindset premise that capability is developable, so the system’s first job is development rather than sorting (Carol S. Dweck, 2006). This chapter builds the replacement in three layers: continuous conversations as the operating rhythm, peer and 360-degree input as the widened evidence base, and the unbundling of development from evaluation and reward as the structural change that makes honesty possible.
9.2 Real-Time Feedback and Continuous Conversations
9.2.1 The Check-In as the Core Ritual
The replacement’s beating heart is the check-in: a short, recurring, forward-looking conversation between employee and manager, typically fortnightly to monthly, structured around near-term priorities, recent work, obstacles, and needs (Peter Cappelli & Anna Tavis, 2016). Deloitte’s redesign made the principle explicit in a rule of thumb: performance conversations should happen at the rhythm of the work, which for project-based teams means weekly or fortnightly, because feedback about this sprint’s work can still change next sprint’s (Marcus Buckingham & Ashley Goodall, 2015). The check-in inverts the appraisal’s tense: where the annual review litigates the past, the check-in steers the future, and its cadence, not any single conversation’s brilliance, is what produces the effect, exactly as sprint reviews outperform post-mortems in Chapter 5.
Real-time feedback extends the rhythm beyond the scheduled conversation: specific, behavioural, near-instant observations, delivered close enough to the event that the recipient can connect feedback to behaviour and act on it. Digital platforms can carry this flow, prompting, recording, and making it visible across a dispersed team (Stefan Strohmeier, 2020), but the tools are carriers, not causes. The GE case of Chapter 4 stands as the caution: an app distributes a feedback culture; it cannot create one.
Continuous feedback fails in two directions. Delivered as a constant stream of evaluation, it becomes surveillance, and people respond by performing safety rather than taking the interpersonal risks that learning requires; Chapter 3’s psychological safety evidence applies with full force, since feedback-seeking is itself a risk behaviour that only safe climates sustain (Amy Edmondson, 1999). Delivered without skill, high frequency multiplies a manager’s bad habits, vague praise, drive-by criticism, at higher volume. Cadence is necessary; safety and capability make it useful.
Every redesign that failed skipped the same investment: conversation capability. Before retiring ratings, train and rehearse the check-in itself, asking before telling, describing observed behaviour rather than characterizing the person, agreeing one commitment per conversation, and equip managers with a light structure such as: priorities since last time, what helped, what hindered, what happens next, what support is needed. Fifteen disciplined minutes a fortnight outperforms ninety improvised ones a year.
9.3 Peer Reviews and 360-Degree Feedback
9.3.1 Widening the Evidence Base
Agile ways of working break the single-rater assumption on which the traditional appraisal rested. When people work in the cross-functional squads of Chapter 3, the manager sees a fraction of anyone’s performance; peers, internal customers, and squad-mates see the rest. Peer review brings squad-mates’ observations into the picture at working rhythm, often as lightweight recognition and retrospective-style reflection within the team. 360-degree feedback formalizes the widened lens: structured input from manager, peers, direct reports, and internal customers, aggregated so the individual can see patterns invisible from any single vantage point (Peter Cappelli & Anna Tavis, 2016).
The design choices determine whether the widened lens illuminates or distorts. Confidentiality must be explicit and honoured; raters must be asked about behaviour they have actually observed, not asked to score abstractions; and volume must be rationed, since survey fatigue degrades every response after the third request in a month. Most consequentially, purpose must be fixed in advance: 360 input used for development is answered honestly; the same instrument wired to pay converts colleagues into constituencies and feedback into campaigning.
| Design choice | Development use | Reward use |
|---|---|---|
| Rater honesty | High: candour helps the recipient | Degrades: ratings inflate or trade favours |
| Recipient stance | Curiosity, pattern-seeking | Defence, source-guessing |
| Team effect | Normalizes mutual feedback | Injects politics into peer relations |
| Appropriate role | Primary use of 360 input | At most one input, heavily moderated, never automatic |
9.4 Unbundling Development, Evaluation, and Reward
9.4.1 Three Decisions, Three Instruments
The traditional appraisal bundled three different jobs into one November meeting: helping the person grow, assessing how they performed, and dividing the pay budget. The bundle is why the meeting satisfied no one; the presence of the pay decision crowded out everything developmental, a dynamic every participant recognizes (Peter Cappelli & Anna Tavis, 2016). Agile designs unbundle. Development runs continuously through check-ins and 360 input. Evaluation, where it survives, becomes lighter and closer to the work: Deloitte’s version replaced year-end ratings with four forward-looking “performance snapshot” questions answered by the team leader at each project’s end, generating frequent, decision-oriented data instead of one annual verdict (Marcus Buckingham & Ashley Goodall, 2015). Reward decisions then draw on this accumulated evidence in a separate calibrated process, at their own cadence.
Unbundling has a price, and honest design acknowledges it. Removing ratings does not remove judgment; it relocates judgment into calibration discussions that are less visible to the employee, which can feel less fair even when it is more accurate. Organizations that retired ratings and later reinstated simplified versions, Microsoft among the pioneers of both moves, mostly did so to restore a legible link between performance and pay. The agile position is not that evaluation is illegitimate but that it must not monopolize the system: the test of a performance design is whether the feedback loop that improves work survives contact with the decisions that price it.
flowchart TD
W["Work, sprint by sprint"] --> CI["Check-ins and real-time feedback<br>(development, continuous)"]
W --> PS["Project-end snapshots<br>(evaluation, frequent and light)"]
CI --> DEV["Growth plans,<br>coaching, mobility"]
PS --> CAL["Calibration<br>(periodic)"] --> PAY["Reward decisions"]
style W fill:#e8eaf6,stroke:#5C6BC0
style CI fill:#e3f2fd,stroke:#1976D2
style PS fill:#fff8e1,stroke:#F9A825
style DEV fill:#e8f5e9,stroke:#388E3C
style PAY fill:#ede7f6,stroke:#7E57C2
9.5 Case Studies
9.5.1 Case Study 1: Adobe, Inventing the Check-In
Adobe’s 2012 abolition of the annual review is the movement’s founding case. The company replaced stack-ranked annual appraisals with “Check-In”: quarterly-or-better conversations on expectations, feedback, and growth, no ratings, no forced distribution, no standard form, with managers given a “Center of Excellence” of resources and coaching rather than a compliance template. Managers received a budget and discretion for compensation decisions informed by their ongoing knowledge of performance. Adobe reported roughly 80,000 manager hours returned to real work and a marked fall in voluntary attrition in subsequent years, alongside an internal finding the movement made famous: the old process had actively manufactured disengagement, with resignations spiking after each rating cycle (Peter Cappelli & Anna Tavis, 2016).
Discussion Questions:
- Adobe removed the form as well as the rating. What does standardized paperwork add to and subtract from a feedback system?
- Compensation discretion moved to managers. What must be true of calibration and data for that discretion to stay fair?
- Attrition fell after the redesign. Construct two rival explanations and the evidence that would separate them.
9.5.2 Case Study 2: Deloitte, Redesigning on Evidence
Deloitte’s redesign began with arithmetic and a validity problem: about two million hours a year spent on reviews, and research showing raters’ scores said as much about the rater as the rated. Its replacement had three parts: weekly project-rhythm check-ins between team leaders and members; four forward-looking snapshot questions per team member at project end, framed as what the leader would do with the person, would always want them on the team, would promote today, rather than abstract trait scores; and separated processes using the accumulated snapshots for talent and pay decisions (Marcus Buckingham & Ashley Goodall, 2015). The design’s premise generalizes: measure frequently, in small honest units, at the rhythm of the work, and let decisions draw on the accumulated stream rather than on one annual reconstruction.
Discussion Questions:
- Why might “would I always want this person on my team?” produce more reliable data than “rate this person’s teamwork 1 to 5”?
- Weekly check-ins at Deloitte’s scale demand enormous manager time in total. Reconcile this with the two-million-hour critique of the old system.
- Snapshot data flows to decisions the employee does not directly see. Design the transparency mechanisms you would add, using Chapter 3’s principles.
9.6 Summary
The annual appraisal declined because its cycle no longer matched the work, its cost outran its validity, and its backward-looking judgment failed to develop anyone (Marcus Buckingham & Ashley Goodall, 2015; Peter Cappelli & Anna Tavis, 2016). The agile replacement runs on three layers. Continuous conversations, check-ins at the rhythm of the work plus specific real-time feedback, form the operating cadence, effective only where psychological safety and manager conversation skill support them (Amy Edmondson, 1999). Peer and 360-degree input widen the evidence base to match cross-functional working, provided purpose, confidentiality, and volume are designed deliberately. Unbundling separates development, evaluation, and reward so the growth conversation survives the pay decision, with light, frequent instruments such as project-end snapshots replacing the annual verdict. Adobe and Deloitte supply the founding designs and their evidence. Chapter 10 adds the goal-setting layer, OKRs, and the agile learning methods that convert feedback into capability.
Annual appraisal · Check-in · Real-time feedback · Cadence · Feedback culture · Peer review · 360-degree feedback · Rater validity · Performance snapshot · Unbundling · Calibration · Forced distribution
Summary
| Concept | Description |
|---|---|
| Why the Appraisal Fell | |
| Decline of the annual appraisal | The movement from about 2012 in which leading firms replaced ratings with frequent conversations |
| Cycle mismatch | A yearly appraisal cycle describing project work that finished months earlier |
| Cost-validity critique | Deloitte's two million review hours set against evidence that ratings measure the rater |
| Continuous Conversations | |
| Check-in | A short, recurring, forward-looking conversation on priorities, obstacles, and needs |
| Rhythm of the work | Matching conversation cadence to the pace of the work rather than the calendar |
| Real-time feedback | Specific, behavioural observations delivered close enough to the event to act on |
| Feedback and psychological safety | The dependence of honest feedback exchange on a climate safe for interpersonal risk |
| Manager conversation capability | The trained skill of asking, describing behaviour, and agreeing one commitment |
| Peer and 360 Input | |
| Peer review | Squad-mates' observations entering the picture at working rhythm |
| 360-degree feedback | Structured input from manager, peers, reports, and customers, aggregated for pattern-seeing |
| Observed-behaviour rule | Asking raters only about behaviour they have actually observed |
| Purpose separation for 360 | Fixing development or reward use in advance, since wiring 360 to pay corrupts candour |
| Unbundled Decisions | |
| Unbundling | Separating development, evaluation, and reward into distinct instruments and cadences |
| Performance snapshot | Deloitte's four forward-looking would-you questions answered at each project's end |
| Calibration | The periodic, evidence-weighing process through which reward decisions are made |
| Legibility trade-off | The fairness-legibility cost of moving judgment from visible ratings to calibration rooms |
| Case Evidence | |
| Adobe Check-In | Adobe's 2012 abolition of ratings for quarterly conversations, returning 80,000 manager hours |
| Deloitte redesign | Deloitte's weekly check-ins, project-end snapshots, and separated decision processes |