How usage is measured: tokens, searches, and what counts
What is a token?
A token is the unit of text we process. As a rough guide, one token is about four characters of English, or about three quarters of a word — so a thousand words is roughly 1,300 tokens. Images, audio, and other rich content are converted to a token count too, at a different rate from plain text.
These are not the same as the tokens your LLM provider bills you for. They are counted separately and they measure a different thing: what Supermemory ingested and indexed, not what your model read or generated.
There is more than one meter
Usage is tracked on separate meters, and it is worth knowing which one your workload draws from:
- Tokens — content processed during ingestion. Plain text and rich content are metered at different rates.
- Search queries — counted per retrieval request, not by the size of the request or the results.
- Operations — platform work outside of ingestion and retrieval, such as certain immediate-processing modes.
SuperRAG workloads are metered on their own token meters, separately from memory ingestion. The full breakdown lives in the Billing and usage docs.
What consumes tokens?
Adding content is what consumes the token meter. That includes:
- Adding a document or a memory directly
- Uploading a file
- Batch ingestion
- Ingesting or updating a conversation
- Anything a connector imports on your behalf
Connector imports are the ones people forget. A connector sync ingests content exactly like a manual upload does, and it counts the same way.
Do searches consume tokens?
No. Searches are metered separately, as queries. Each retrieval request counts as one query regardless of how much text comes back, so a search that returns twenty results costs the same as one that returns two.
Profile calls that run retrieval count against the search meter as well. Calls that don't retrieve anything — listing your documents, checking processing status, reading settings — are not searches and don't increment that meter.
The practical consequence: a read-heavy application and a write-heavy application produce very different-looking usage. If your token usage is flat but your bill moved, look at the search meter.
Re-adding the same content doesn't cost twice
Content is deduplicated. When you re-add something under the same customId, we compare the new version against the previous one and bill only the net-new difference. Unchanged content is fully discounted.
This means a connector re-sync, a repeated upload, or pushing a growing conversation history repeatedly does not redraw your balance for the parts that haven't changed. To get this discount, keep the customId stable, stay within the same organization and key, and update the document rather than deleting and recreating it.
When does usage reset?
Included usage resets at the start of each billing period and does not roll over. Purchased top-up credits are not tied to the period — they remain until you use them.
Where can I see my current usage?
The billing and usage page in the Console shows live usage for the current period, broken out per meter. If you want to pull the same numbers programmatically — for your own dashboards or alerting — there are usage endpoints documented in the Billing and usage docs.
Check the Console first when usage looks wrong. It tells you which meter moved, which narrows the cause immediately.
Why did my usage spike unexpectedly?
Almost always one of these:
- A first connector sync. Connecting a large Drive, Notion workspace, or mailbox imports the backlog, not just new content going forward. The initial import is usually the largest single ingestion an account ever does. Scope the connector to the folders or labels you actually need before the first sync.
- An agent loop rewriting memories. An agent that writes back after every turn can ingest continuously. Deduplication absorbs the unchanged parts, but content that genuinely changes each time is genuinely new content.
- Recreating instead of updating. Deleting and re-adding a document is billed as new content. Updating it in place under the same
customIdis not. - Search volume, not content volume. An agent that searches several times per user turn multiplies query usage quickly, even when nothing is being ingested.
- Rich content. Images, PDFs with scanned pages, and audio meter differently from plain text, so a small number of files can carry a large token count.
What happens when I hit my limit?
Metered API calls can start being rejected — typically with an HTTP 402 — depending on your plan and overage settings, until you top up, enable overages, or the billing period resets.
For the full picture on overages, top-ups, and what to do when you're blocked, see the “Understanding token limits and overages” article in this help center.