Claude Code
Claude Code adoption metrics: measure team value without ranking engineers
Build a team-level Claude Code scorecard with clear denominators, delivery measures, privacy rules and a worked decision about rollout expansion.
September 15, 2026/9 min read/Claude Certified Engineers

Your Claude Code dashboard shows more activity, but the engineering review still ends with the same question: should you expand access, change the training, or fix the delivery process? A list of the heaviest users cannot answer it. Measure adoption at team level, reconcile what each source counts, and pair usage with the outcomes that would justify your next decision.
That approach also makes capability verification more useful. An engineer can generate substantial output and still need help reviewing it. Another can use Claude selectively while catching a serious defect. A team's readiness rests on whether people can explain, test, and own the changes they ship. A usage chart is one input to that assessment.
Set the decision before choosing the denominator
Start with a sentence such as: “At the next monthly review, we will decide whether to expand the approved workflow to another service.” This gives the scorecard an owner and a consequence. A license-reallocation decision needs a different view from an assessment of whether a team can safely take over a production system.
Define the eligible population before opening an activity chart. For an adoption review, it might be engineers assigned to the pilot service who have approved access and suitable work during the measurement period. Record exclusions, including extended leave and roles outside the workflow. Keep the definition stable through the comparison; report a changed population as a new cohort.
The distinction matters even when two reports use the same label. DX lets readers choose between team-assigned contributors and license holders as the population. Its active-user calculation counts any positive metric in the usage report. DX also documents why its daily count can differ from the Anthropic Console's session-start count when a session spans days. Those are different definitions of activity, not necessarily a broken integration. DX's adoption documentation describes both the denominator and the activity rule.
Name the rule in your report: “weekly users with recorded activity / eligible pilot engineers.” Do not average daily adoption percentages into a weekly rate. Deduplicate people across the week, then divide by the eligible population for that week.
Use a small metric dictionary that survives a review
The following dictionary is a proposed operating model. Each row answers a different question; the rows should remain separate rather than collapse into a single productivity score.
| Measure | Definition and source | Decision it supports | Interpretation limit |
|---|---|---|---|
| Weekly reach | Unique eligible people with recorded activity divided by eligible people; approved usage export joined to the access roster | Find access or workflow friction | A recorded event does not establish useful work |
| Repeated use | Eligible people active in at least three of the last four complete weeks divided by the stable eligible cohort; deduplicated weekly exports | Decide where to offer workflow coaching | Three weeks is a proposed consistency rule, not a capability threshold |
| Assisted contribution | Merged PRs attributed to Claude Code divided by all merged PRs in the same covered repositories; contribution export | Understand where assistance reaches shipped code | Attribution does not establish quality or time saved |
| Delivery flow | Median and upper-quartile elapsed time from first commit to production for the same service and change class; version control plus deployment records | Locate waiting or review bottlenecks | A changed mix of work can move the distribution |
| Release safety | Deployments requiring immediate intervention divided by deployments, with the underlying counts; deployment and incident records | Decide whether expansion needs to pause | Small samples are unstable; one serious incident needs investigation regardless of the rate |
| Operating cost | Actual billed usage plus allocated license charges and separately reported review effort; finance records and team estimates | Assess whether the workflow is affordable | Token cost alone omits licenses, review and rework |
| Demonstrated capability | Results from a small agreed work sample: explain a diff, detect a seeded issue, test a boundary and describe recovery | Choose the next enablement exercise | A sample supports a training decision, not a claim about every future task |
For every row, save the aggregation level, cohort, reporting window, decision owner and privacy rule alongside the value. Use one service for delivery and release measures, and the corresponding team cohort for adoption. Avoid combining unrelated services just because they share a manager.
DORA defines change lead time from commit to production and change fail rate in terms of deployments requiring immediate intervention. It recommends interpreting delivery measures in an application's context and warns against relying on a single metric. Those principles support keeping the flow and safety rows together. DORA's measurement guide supplies the definitions; the coaching and reporting rules here are recommendations.
The capability row deserves a named reviewer. Ask an experienced engineer to examine how participants reason about a change, including where Claude's output needs correction. Reuse the same task boundaries and rubric across the cohort. The CI evaluation guidance can inform the test boundaries. Keep the learning review separate from hiring or compensation decisions.
Reconcile assistance, acceptance and delivery
Keep accepted edits separate from merged contributions. Anthropic's analytics documentation says accepted lines do not track subsequent deletions. Its contribution metrics identify merged PRs containing Claude-assisted code, with conservative attribution rules and exclusions. A PR can qualify with a small assisted contribution; that does not mean Claude produced the whole change. Anthropic's analytics reference explains the distinction.
This is why “assisted PRs increased” cannot stand in for “delivery improved.” The first is an attribution statement. The second requires evidence about the delivery process and its results. Neither establishes how much faster the same work would have happened without the tool.
Before comparing periods, check repository coverage, user identity mapping, date boundaries and data freshness. Keep an explicit “unknown” category for missing coverage. Do not convert missing activity into zero adoption, and do not add figures from overlapping exports without deduplicating their identities and events.
If you use OpenTelemetry, keep the collection policy narrow. Anthropic documents optional controls for logging prompt content and tool details, as well as usage metrics. Enabling content collection expands what the telemetry can reveal. An adoption review usually needs aggregated activity and cost, not the contents of an engineer's conversation. Review the monitoring controls before deciding what to retain.
Work through a decision with the whole team in view
Consider this hypothetical pilot for one service. Twelve engineers are eligible in both four-week windows. Repository coverage, reporting rules and the class of changes are held constant. The numbers below illustrate a review; they are not customer results or performance targets.
| Observation | Earlier window | Later window |
|---|---|---|
| Active people in at least three of four weeks | 6 of 12, or 50% | 9 of 12, or 75% |
| Merged PRs with attributed assistance | 18 of 40, or 45% | 36 of 48, or 75% |
| Median first-commit-to-production time | 3.0 days | 2.4 days |
| Upper-quartile first-commit-to-production time | 5.0 days | 6.0 days |
| Deployments requiring immediate intervention | 1 of 20, or 5% | 2 of 20, or 10% |
Repeated use has risen by 25 percentage points. The median lead time fell by 20%, calculated as (3.0 - 2.4) / 3.0. But the upper quartile worsened, and two deployments required intervention. With only twenty deployments per window, the change in failure rate is too thin a basis for a broad claim about reliability. It is enough to require examining the incidents before expansion.
Suppose that examination finds both interventions involved untested migration boundaries, while the slowest changes waited for a specialist reviewer. The next action is specific: keep the cohort stable, add a migration-boundary exercise, and arrange reviewer coverage. Do not respond by asking everyone to generate more code.
At the following review, inspect whether the agreed boundary tests were added and whether review waiting time changed. Preserve incident severity and impact in the record; a lower count cannot erase a more damaging failure. If the two incidents turn out to have unrelated causes, record that too. The scorecard should support investigation rather than supply a convenient story about Claude.
A before-and-after comparison still cannot isolate the tool's effect from staffing, task selection or process changes. For a stronger evaluation, agree comparable work and a comparison design before rollout. For a routine operating decision, it may be enough to identify the next bottleneck without claiming a causal productivity gain.
Make the privacy boundary operational
Write down who sees the raw exports and who sees the team summary. A workable starting policy is to keep identifiable usage with the platform administrator, give the engineering review aggregated figures, and prohibit use of those figures for individual performance rankings. Let engineers correct identity or access errors before a report drives action.
Choose a minimum cohort size with your privacy reviewer. For example, you might suppress groups smaller than five and disallow filters that isolate one person. Five is an illustrative starting point, not a guarantee of anonymity: a known on-call schedule or a unique specialist role can identify someone in a larger group. Combine groups only where the resulting measure still makes sense; otherwise withhold the breakdown.
Set a retention period that fits the decision window and the organization's policy. A dashboard permission is not a reason to retain identifiable records indefinitely. For license follow-up, an administrator can contact the affected person privately and check leave, access or workload before reallocating anything.
The same separation applies to certification. Keep credential records and assessed work samples in their appropriate process. Do not infer credential quality from token consumption or label low activity as a failed skill assessment. A capability review should make clear what was observed and what remains untested.
End the review with a decision someone owns
Close each review with a compact record: the service and cohort, the reporting window, the definitions used, the observed change, the explanation still being tested, and the action with an owner and review date. Include what would reverse the decision. This is the useful output of the scorecard.
For the hypothetical pilot, that record would hold expansion until the migration exercise and reviewer coverage are checked. For another team, stable delivery and verified work samples might justify a limited expansion. The enterprise rollout checklist covers the surrounding access and governance controls; this review supplies evidence for the next step.
Before your next meeting, pick one service and write the denominator and activity rule at the top of its report. If nobody can explain what a changed number will make the team do differently, leave that number off the decision page.