NearSync Help

AI

Watching AI Usage Six numbers, what each one is really telling you, and what the counters do not cover.

Usage is where you see what the AI is doing. Six numbers, each answering a different question.

The six

Number Answers
Credits used What it is costing you
Average latency Whether it feels fast
Total prompts Whether anybody is using it
Tokens consumed How much work is going through
Active models How many are actually in play
Error rate Whether it is working

Each shows how it has moved, which is more useful than the value. A five percent error rate that has been five percent all quarter is a known quantity. One that was zero last week is an incident.

The two to watch

Error rate, because a rising one is the only number here that means something is broken. Everything else is information; this is an alarm.

Errors show up to users as an assistant that did not answer, and people do not report that. They try once more and stop using it. By the time anybody mentions it, the number has been elevated for weeks.

Average latency, because it silently decides whether the AI gets used at all.

There is a threshold, somewhere around three seconds, past which people stop asking. They do not decide to; they simply stop finding it worth the wait. If latency has crept up after a model change, that is the finding, and Choosing your model is where it gets fixed.

Total prompts is an adoption number

Not a cost number. It answers whether the AI is part of how people work.

Flat and low means it was set up and nobody uses it, which is nearly always a knowledge problem: people tried it, got generic answers, and concluded it was not for them. Teaching the assistant is the fix, not encouragement.

Rising steadily is what adoption looks like. There is no target; the right number is whatever your team finds useful.

Spiking then collapsing is the pattern after a launch, and it means the first impression was poor.

Credits

What it costs, in the unit your plan is written in. What a credit is and what spends one is covered in Credits and what uses them.

Watch the trend, not the total. A total is only meaningful against your allowance; a trend tells you whether something changed.

A sudden rise usually has one cause, and it is usually an agent. An agent scheduled too often, or with conditions loose enough to match constantly, spends steadily and quietly. Check agents before anything else.

Active models

How many models are genuinely being used.

One is a warning. It means everything depends on a single provider, and there is no fallback in practice regardless of what is configured. Worth checking against How it is set up.

Tokens, briefly

The unit of work models are measured in, roughly corresponding to pieces of words.

You do not need to think in tokens. The number is useful for one thing: compared against prompts, it tells you whether answers are getting longer. A steady climb in tokens with flat prompt numbers means the assistant is saying more per question, which is usually a persona or model change rather than anything people did.

What the counters do not cover

Worth being clear about, because it changes how you read every number above.

These figures count assistant conversations. Somebody asks, a model answers, it is counted.

Not everything the AI does is a conversation. Agent runs, speaking out loud, meeting transcription and the indexing behind Knowledge are AI work that does not all arrive in the prompt count.

So treat the prompt figure as a floor on activity, not a total. It is an accurate picture of how much people are asking, and an incomplete picture of everything happening. For what your plan is actually consuming, the credits article is the reference.

Reading two numbers together

Individually the six are informative. In pairs they diagnose.

Prompts up, credits flat means people are asking more and the answers are getting shorter or cheaper. Usually good.

Credits up, prompts flat means something is consuming without anybody asking. Agents, almost always.

Latency up, error rate up together points at a provider having a bad time rather than at anything you changed.

Latency up, error rate flat points at a model change. Something is answering as reliably as before and taking longer to do it.

Everything flat and prompts falling is the quiet failure: nothing is broken, and people have stopped bothering. That is the hardest one to spot and the most worth catching, because nobody reports it.

Using it monthly

Four questions, ten minutes.

Is the error rate where it was? If not, check providers and models.

Has latency crept? If so, something changed, usually a model.

Is usage growing or dying? Dying means knowledge, not training.

Did credits jump? If so, look at agents first.

When to look

Monthly is enough, plus whenever somebody complains.

The complaint case is the important one. "The AI has got worse" is almost never imagined, and it is almost never about the answers being wrong. Nine times in ten it is latency, and this page will show you that in five seconds where a conversation about answer quality would take an afternoon and settle nothing.

Share it once a quarter

The trend, not the totals, and to whoever sponsored the AI. A rising usage line is the clearest evidence it was worth doing, and a falling one is a problem best raised early.

What it will not tell you

Whether the answers were any good. Nothing here measures quality. A model answering fast, cheaply and wrongly looks excellent on this page.

Who asked what. This is operational monitoring, not a record of individual conversations.

Which questions failed to help. The most valuable thing you could know, and not available. The nearest substitute is asking people, which is worth doing once a quarter.

5 minUpdated 28 July 2026

Did this answer your question?

No, ask a person