Pheebs is an open-source telemetry tool built by Eversynced to understand how developers work with AI coding agents and what the models they run are costing them. It installs quietly inside the AI coding agents Claude Code, Cursor, and Codex through hooks, capturing lightweight interaction signals: the shape of the session, not its contents. Hooks and OpenTelemetry go in; honest proficiency reads come out. The product is built for teams that want an evidence-based answer to a simple question — how is AI coding actually being used here, and what is it costing?
AI coding agents are fast and their output often looks polished, which makes them very hard to assess by feel. Polished output can hide missing verification. Over-provisioned models can burn budget without anyone noticing. Follow-up prompts spent repairing AI-generated breakage can look indistinguishable from healthy iteration unless someone measures them. The site frames this through a set of observations: model spend that buys nothing, where thousands of dollars of last month's model spend went to a bigger model than the work needed; AI code that ships unchallenged, where a majority of AI-written lines in a payments service shipped with no check; AI edits that never had a test, typecheck, or build run behind them; rework hiding inside the speedup, where follow-up prompts were fixing something the AI broke rather than moving the work forward; enablement skills that either caught on weekly or never caught on at all; and teams that never run tests inside the agent loop at all. On that last point the site is explicit — that is a missing harness, not a skills gap, and Pheebs is positioned to help teams tell which situation applies to them.
Pheebs works with three coding agents: Claude Code, Cursor, and Codex. The client sits inside each agent via hooks, and the coverage spans 17 event types, from session_started through artifact_found. Events include session starts and ends, prompts, skill and slash-command expansions, sub-agent spawns, tool calls and failures, compaction, and background tasks. A sample Claude Code stream shows the granularity in practice: session_started with a codebase and model, prompt_submitted with a prompt length and intent label, tool_use_completed entries for an Edit and a Bash test run, context_compacted with a trigger type, and turn_ended with a background task count. Cursor connects through hooks, while Claude Code and Codex connect through hooks plus OpenTelemetry.
Every field Pheebs records is deliberately lightweight, and the tool is explicit about what it never captures. Source code and file contents are never stored. File paths and directory structures are excluded, with a repository recorded only as org/repo from the git remote. Prompt text is never stored — a prompt becomes a character count, with an intent label added when the prompt intent classifier is enabled. Command strings are read in process, so npm test is recorded as tool_intent: test_run rather than as text. Names and emails are avoided: a developer is the id behind their Pheebs token, stamped by the backend, or a truncated hash of their git email when no token is set, and a GitHub handle is never looked up. The site sums it up bluntly: no code, no file paths, no stored prompt text — the shape of the session, never its contents.
The capture pipeline is documented step by step. First, a hook fires. Second, lightweight fields are extracted: event type, durations, counts, models, and trigger types, with a prompt reduced to a character count and, when the prompt intent classifier is enabled, an intent label — the text itself is never stored. Third, identity and repo are resolved, using the developer id behind the token or a truncated hash of the git email, and the codebase as org/repo from the git remote. Fourth, every event is stored in a local JSONL log, and with a token set it also goes to the backend. Fifth, OpenTelemetry rides along: Claude Code and Codex export native OTel metrics and logs through the Pheebs proxy. The client offers four routes to any backend — self-hosted, or managed by Eversynced — and the local JSONL stays the durable copy either way. Configuration is deliberately minimal: set a base-url and set a token. Both need to be set or nothing is posted, and unsetting either one stops sending.
The backend contract is documented, with a reference backend available in the Pheebs repo. POST /ingest carries one event envelope per request. POST /validate-token resolves a token to an identity and its consent flags. POST /classify-prompt takes one prompt in and returns one label, and it is the only route that receives raw text. POST /otel/v1/{signal} is an OTLP passthrough, so no observability credential ever ships in the client. GET /insights is optional and covers what one developer can see about their own work.
On top of the raw events, Pheebs renders a proficiency model organized into six competency areas. Models covers which models are in play: model choice, effort settings, plan mode, and autonomy modes. Artifacts covers the reusable config that shapes the agent: skills, sub-agents, slash commands, and context files. MCP covers live connections to external systems such as tickets, databases, browsers, and documentation. Evals covers verification wired into the agent loop: tests, typecheck, lint, build, and review passes. Context management covers deliberate use of the context window, including compaction and the save, resume, and clear lifecycle. Orchestration covers more than one agent at a time: sub-agents, parallel work, worktrees, hooks, and plugins.
Each competency is tracked in one of three states. Unobserved means the practice never showed up in the window. Adopted means it showed up at least once. Recurring means it showed up in at least three of the last four active weeks. The coverage index summarizes this per engineer as the share of applicable practices at Recurring. Alongside the competencies sit five judgement signals, split between output side and input side. On the output side, verification coverage is the share of AI edits followed by a verification action — a test run, typecheck, lint, build, or a check against a spec. Pushback rate measures how often the engineer challenges AI output instead of accepting it, a signal the site notes collapses exactly when output looks polished. Refinement-to-repair ratio distinguishes whether follow-up prompts refine intent (healthy iteration) or repair breakage (rework). Wholesale-accept rate captures sessions with no pushback, no repair, and no verification, weighted by lines changed — described as the composite red flag of polished output with no questions asked. On the input side, model-fit rate is the share of sessions whose model class matched the size of the work.
Model-fit is the one signal with a price attached. A reporting view shows savings opportunity against list-price spend, contrasting the models used with the work as sized, and it carries a coverage breakdown — complete, incomplete, no telemetry, unpriced — because decisions and figures come from complete sessions only. Three principles govern the approach. Tasks are sized: every task prompt gets a scope, from a one-file change to open-ended design, and a session is judged on its hardest prompt. Misses count both ways: an over-provisioned session burns budget silently, while an under-powered one shows up as repair prompts. And Pheebs is an audit, not a router: it never intercepts a prompt or switches a model on anyone's behalf — it reads the gap and prices it, and the decision stays with the team.
Reporting built on top of Pheebs renders the model in several views. A practice adoption funnel shows one bar per competency, split by how many engineers have not acted on it, acted once, or acted week after week, with Unobserved and Adopted flagged as the competencies to be intentional about. A practice heatmap puts every engineer against every competency; a cold column means the team is missing the setup and practice for that competency, which is described as a structural fix, while a cold row calls more strongly for coaching. A per-engineer view shows how much of each competency has become habit and sums it up in a coverage index that can be tracked over time. A signals-by-engineer table lists verification, pushback, refine-to-repair, wholesale accept, and model-fit per person alongside a team median. The guidance is direct: one weak number is a coaching conversation, but a weak column across the whole team is a structural gap. The sample views on the site are labeled illustrative data.
Pheebs can be deployed in two ways. Self-hosted means you stand up the backend and telemetry goes from your developers' machines to your own infrastructure — Eversynced never sees it. That option includes the full client with all three agents under Apache-2.0, a documented contract and a reference backend in the repo, raw JSONL you can query with whatever you already use, and no account, no key, and no requests from Eversynced. Managed means Eversynced runs it, along with the reporting on top: the same open-source client pointed at an operated backend, with the proficiency model rendered as reports and dashboards. That is the AI Enablement Assessment service — a 30-day telemetry sprint that ends in an executive debrief and a plan for the gaps, including the model-fit gap priced in dollars from the team's real sessions, with insights tracked over time.
Installation is a single npm command, followed by pheebs init for interactive setup across all three agents and pheebs doctor to check the wiring. In practice the product serves teams that want to see where AI budget actually goes, teams diagnosing whether weak AI results are a setup problem or a coaching problem, and individual developers who want their own honest read on their practice. Eversynced runs Pheebs on itself: every Eversynced engineer is instrumented with it, and it powers the measurement layer of the company's AI delivery framework, which is the same reporting that ships with the AI Enablement Assessment run for client teams.
The takeaway is that Pheebs turns an otherwise invisible activity — how a team works with AI coding agents and what those agents cost — into measured, priced evidence. It does so without storing the work itself, and it leaves every decision with the team: an audit rather than a gatekeeper.