Skip to content
AI Performance Insights

See how well your team uses AI, not just how much

AI Performance Insights turns every captured session into 16 signals across efficiency, engagement, quality and sessions, so you know who to coach, what to fix and which habits to spread.

Four groups, one view

Four questions every engineering lead asks about AI

Each group answers one question with a handful of focused signals. Start from the headline number, then open the group to see the developers behind it.

Efficiency

Find the habits that waste tokens

Six signals show where context is resent, prompts are oversized or the AI loops through too many tool calls. Pick a signal to see who it lists, how to read it and what to do about it.

Cache efficiency

Developers with the lowest cache hit rate, lowest first.

  • Kai84%

    98.0K input tokens

  • Priya91%

    184.2K input tokens

  • Nina94%

    135.2K input tokens

  • Lena94%

    132.6K input tokens

  • Ben95%

    107.0K input tokens

How to read it

A low rate means the same context is sent again and again at full price, instead of being read back from the cache for a fraction of it.

What to do about it

Keep long-running work in one session and keep project context in CLAUDE.md, so it is cached once and reused.

Engagement

Spot quiet seats before renewal

See who barely uses AI, who is sliding and whose prompts burn the most tokens, so you can step in early instead of finding out at renewal.

  • Low engagement

    Developers contributing less than 10% of team output over the last 30 days.

    Everyone is actively engaged

    Every developer contributes at least 10% of team output

    What to do: Check in before the next renewal: a short onboarding session often turns a quiet seat into an active one.

  • Least active

    The lowest token users, ranked from the bottom, leaving out the top 5.

    • #9
      Kai480.6K
    • #8
      Ben569.0K
    • #7
      Lena594.6K
    • #6
      Nina613.9K

    What to do: Look at the trend for each name, and ask whether something is in the way, such as access, setup or the kind of work they do.

  • Prompt burn

    Highest average tokens per prompt, highest first. At least 3 prompts.

    • Priya81.0K / prompt

      21 prompts · 1.7M tokens

    • Maya63.4K / prompt

      13 prompts · 824.4K tokens

    • Alex59.7K / prompt

      15 prompts · 895.2K tokens

    • Raj52.4K / prompt

      13 prompts · 681.6K tokens

    • Sam51.9K / prompt

      15 prompts · 777.9K tokens

    What to do: Break big asks into smaller steps, and turn on Prompt Enhancement to add the right context instead of all of it.

Quality

Check the right model is doing the work

See who leans on AI the most, how much the AI writes back and who sends routine work to the most expensive model.

  • Top performers

    The 5 developers who used the most tokens in the selected period.

    • 1st
      Priya1.7M
    • 2nd
      Alex895.2K
    • 3rd
      Maya824.4K
    • 4th
      Sam777.9K
    • 5th
      Raj681.6K

    What to do: Ask your heaviest users to share the prompts and skills that work for them, so the rest of the team can reuse them.

  • Output verbosity

    Developers with the highest ratio of output to input tokens, highest first.

    • Priya0.22× output

      184.2K in · 40.2K out

    • Maya0.21× output

      176.9K in · 36.9K out

    • Alex0.17× output

      203.7K in · 34.1K out

    • Sam0.17× output

      174.4K in · 29.6K out

    • Raj0.17× output

      153.1K in · 26.1K out

    What to do: Ask for diffs or short answers where full rewrites are not needed, and check the output style in team skills.

  • Model spend mix

    Developers with the highest share of tokens on Opus, highest first.

    • Priya62% Opus

      Opus 62%, Sonnet 33%, Haiku 5%

    • Raj48% Opus

      Opus 48%, Sonnet 47%, Haiku 5%

    • Alex31% Opus

      Opus 31%, Sonnet 64%, Haiku 5%

    • Lena22% Opus

      Opus 22%, Sonnet 70%, Haiku 8%

    • Sam12% Opus

      Opus 12%, Sonnet 80%, Haiku 8%

    • Opus
    • Sonnet
    • Haiku

    What to do: Suggest Sonnet for routine edits, tests and small fixes, and keep Opus for the hard problems.

Sessions

See which sessions reach a finished result

Runaway sessions, shallow ones, sessions given up early and sessions that end with a summary, for every developer.

  • Outlier sessions

    Developers whose sessions use more than 2× the team average of tokens. Team average: 152.4K per session. At least 3 sessions.

    • Priya340K / session

      2.2× the team average · 5 sessions · 1.7M tokens

    What to do: Open the session in Activity Tracking to see what happened, then fix the prompt or the context that caused it.

  • Shallow sessions

    Developers with the fewest prompts per session, lowest first. At least 3 sessions.

    • Priya4.2 / session

      5 sessions · 21 prompts

    • Maya4.3 / session

      3 sessions · 13 prompts

    • Sam5.0 / session

      3 sessions · 15 prompts

    • Alex5.0 / session

      3 sessions · 15 prompts

    • Lena5.3 / session

      3 sessions · 16 prompts

    What to do: Pair these developers with someone who runs deep sessions, or start them with a skill that walks through a full task.

  • Session abandonment

    Share of sessions with 2 prompts or fewer and 3 observations or fewer. Sessions with a lot of tool activity are not counted. At least 3 sessions.

    • Priya40%

      2 of 5 sessions abandoned

    • Kai33%

      1 of 3 sessions abandoned

    • Alex0%

      0 of 3 sessions abandoned

    • Sam0%

      0 of 3 sessions abandoned

    • Lena0%

      0 of 3 sessions abandoned

    What to do: Find out what makes people give up: a slow start, missing context or the wrong model, then fix that first.

  • Session completion

    Share of sessions that reached a summary, lowest first. At least 3 sessions.

    • Priya40%

      2 of 5 sessions completed

    • Kai67%

      2 of 3 sessions completed

    • Alex100%

      3 of 3 sessions completed

    • Sam100%

      3 of 3 sessions completed

    • Lena100%

      3 of 3 sessions completed

    What to do: Encourage developers to let sessions finish, so every decision and change is summarized and saved to the Knowledge Base.

Built for coaching

Signals you can act on, not a scoreboard

Every signal is designed to start a useful conversation with a developer, never to rank people on a single number.

  • Minimums before rankings

    Signals that depend on averages only list developers with at least 3 prompts or sessions, so one odd day never puts someone on a list.

  • Every signal explains itself

    Each card says what it measures and what a high or low value usually means, so a lead knows what to look at next.

  • Quiet when all is well

    When nobody crosses a line, the card says so in plain words instead of showing an empty chart.

  • Same filters as the rest

    Filter every signal by developer, client and date range, exactly as you do in AI Insights and the AI Cost Center.

FAQ

AI Performance Insights questions, answered

What each signal measures and how to use it.

16 signals in four groups. Efficiency shows whether tokens are spent well, Engagement shows who is and is not using AI, Quality shows whether the right model is doing the work, and Sessions shows whether sessions reach a finished result.

No. Each signal is a prompt for a conversation, not a score. Top performers, for example, lists who used the most tokens, which shows who leans on AI the most, not who does the best work. Read the signals together and use them for coaching.

Averages are misleading on tiny samples. A developer with one huge prompt would top the input per prompt list, so signals built on averages only include people with at least 3 prompts or sessions, and tool utilization needs at least 5 prompts.

Cached context is read back for a fraction of the normal input price. A low hit rate means the same context is being sent fresh again and again, often because work is spread over many short sessions or project context is pasted in by hand.

A session is abandoned when it has 2 prompts or fewer and 3 observations or fewer. Sessions with a lot of tool activity are not counted, because real work happened in them. A session is completed when it reaches a summary.

Model spend mix shows who sends most of their tokens to Opus. Routine edits, tests and small fixes usually work just as well on Sonnet at a lower cost, while Opus is worth it for hard design and debugging problems.

Yes. Every signal follows the developer, client and date range filters. Low engagement is the one exception: it always looks at the last 30 days, so a short date range never hides a quiet seat.

From the same sessions Activity Tracking captures on every developer machine. Prompts, model calls, tool runs, tokens and summaries are already recorded, so the signals fill in on their own with nothing extra to set up.

Help every developer get more from AI.

Install the agent once per laptop and your performance signals fill in from the sessions your team already runs.

  • Free for up to 5 developers
  • 10-minute setup
  • Your Anthropic key, your spend