The short version
Tokens, sessions, and spend describe consumption. They do not prove that useful work happened. A stronger operating view connects AI cost to a concrete unit of completed work, then checks whether that work improved speed, quality, capacity, or another decision-relevant outcome.
What the source adds
Anthropic's enterprise controls place costs beside practical activity such as artifacts created, files edited, and skills or connectors used. The same product announcement describes group and user views, model controls, spend alerts, and value-oriented Claude Code reporting. The design signal is clear: cost governance works better when financial visibility and operational evidence appear together.
The three-layer measurement chain
1. Consumption
Track active users, sessions, models, and total spend. These numbers are necessary for budgeting and anomaly detection, but they are only the first layer.
2. Useful work
Define the unit that the workflow is supposed to complete: a contract reviewed, a research brief delivered, a reconciliation prepared, or a support case resolved. A useful unit should be countable and should represent finished work rather than activity inside the tool.
3. Outcome
Pair that unit with the result the organization actually values. Examples include turnaround time, rework rate, analyst capacity, escalation accuracy, or decision speed. The outcome prevents teams from celebrating cheaper output that is less useful or less reliable.
A practical value card
For each recurring workflow, report five fields:
- Total AI and connected-tool cost.
- Number of completed useful units.
- Cost per completed unit.
- One quality or business outcome.
- One exception or escalation signal.
This card creates a better management conversation. A rising bill may be healthy when throughput and quality rise faster. A falling cost may be unhealthy when work is abandoned or human rework increases.
Example: contract review
A weak report says that a team used 1,900 sessions. A useful report says that the workflow completed 1,180 first-pass reviews, states the cost per completed review, compares median turnaround time with the prior process, and shows the share routed to qualified human review.
The second report supports a decision: scale, refine, reroute, or stop the workflow.
Implementation cautions
- Do not treat activity as value.
- Do not compare workflows using cost alone when their difficulty and quality requirements differ.
- Keep formulas visible so stakeholders can challenge assumptions.
- Segment results by workflow or team before drawing conclusions from organization-wide averages.
- Use exceptions and failure signals alongside productivity estimates.
Try it
Choose one AI workflow. Replace its most obvious activity metric with one useful unit, one cost-per-unit measure, one outcome, and one exception signal. If the resulting card would not help a manager make a decision, refine the unit until it does.