Pickaxe Learn

Changelog
ImprovementBuild ·

Latency, broken into four numbers

Open workspaces

Choose a workspace, then open Activity and select a conversation.

Latency, broken into four numbers

What it does

Message Insights splits response time into Pickaxe latency, model time-to-first-token, total time-to-first-token, and generation time.

Before this

An expansion. Response time was already reported, as one number that told you an agent was slow without telling you which part of it was.

Why it matters

One aggregate number tells you an agent is slow. These four tell you whether to trim the knowledge budget, switch models, or shorten the output.

How it works

One number tells you an agent feels slow. These four tell you which part to fix:

Pickaxe latency
Our own overhead before the model is called — retrieval, assembly, routing. High here means the knowledge budget or the prompt is doing too much work.
Model time-to-first-token
How long the provider took to start answering. High here is a model or provider problem, not yours.
Total time-to-first-token
What the user actually waits before words appear — the two above combined.
Generation time
How long the answer took to finish once it started. High here means the output is long, and shortening it is the fix.