Pricing
Pinecone Assistant usage is billed monthly. Costs can include:- Minimum usage (Builder, Standard, and Enterprise plans)
- Ingestion (file uploads)
- Tokens (chat, context retrieval, and evaluation)
- Storage
Minimum usage
The Builder, Standard, and Enterprise pricing plans include a monthly minimum usage commitment:
On the Builder plan, the monthly minimum is a flat fee that covers included usage; additional usage beyond Builder limits is blocked rather than billed. On the Standard and Enterprise plans, customers are charged for what they use each month beyond the monthly minimum.
The minimum is a commitment you grow into rather than an extra charge. Once your usage exceeds the minimum, you pay only for what you use.
Examples
Usage below monthly minimum
Usage below monthly minimum
- You are on the Standard plan.
- Your usage for the month of August amounts to $20.
- Your usage is below the $50 monthly minimum, so your total for the month is $50.
Usage exceeds monthly minimum
Usage exceeds monthly minimum
- You are on the Standard plan.
- Your usage for the month of August amounts to $100.
- Your usage exceeds the $50 monthly minimum, so your total for the month is $100.
Ingestion
When you upload or replace files for an assistant, usage is measured in ingestion units. One ingestion unit is approximately 400 tokens (~300 words); exact counts can vary by document.
Multimodal PDF processing uses the same ingestion unit; it is billed at about twice the standard per-unit rate. For current rates, see Pricing.
Multimodal ingestion applies to content processed through the multimodal PDF path. Standard ingestion applies to other supported file types.
Usage and invoices reflect a single ingestion usage line item. With API version
2026-04 or later, a completed file-ingestion operation may include ingestion_units. Use Describe an operation or Track file operations for details.
Tokens
For paid plans, you are charged for the number of tokens used by each assistant. Ingestion is billed separately from chat and context retrieval tokens.Chat tokens
Chatting with an assistant involves both input and output tokens:- Input tokens are based on the messages sent to the assistant and the context snippets retrieved from the assistant and sent to a model. Messages sent to the assistant can include messages from the chat history in addition to the newest message.
- Output tokens are based on the answer from the model.
*1,000,000 input tokens/month to explore Marketplace apps until June 30, 2026.
Chat input tokens appear as āAssistants Input Tokensā on invoices and
prompt_tokens in API responses. Chat output tokens appear as āAssistants Output Tokensā on invoices and completion_tokens in API responses.Context tokens
When you retrieve context snippets, tokens are based on the messages sent to the assistant and the context snippets retrieved from the assistant. Messages sent to the assistant can include messages from the chat history in addition to the newest message.Context retrieval tokens appear as Assistants Context Tokens Processed on invoices and
prompt_tokens in API responses. In API responses, completion_tokens will always be 0 because, unlike for chat, there is no answer from a model.Evaluation tokens
Evaluating responses involves both input and output tokens:- Input tokens are based on two requests to a model: The first request contains a question, answer, and ground truth answer, and the second request contains the same details plus generated facts returned by the model for the first request.
- Output tokens are based on two responses from a model: The first response contains generated facts, and the second response contains evaluation metrics.
Evaluation input tokens appear as Assistants Evaluation Tokens Processed on invoices and
prompt_tokens in API responses. Evaluation output tokens appear as Assistants Evaluation Tokens Out on invoices and completion_tokens in API responses.Storage
For paid plans, you are charged for the size of each assistant.Limits
Pinecone Assistant limits vary based on subscription plan.Object limits
Object limits are restrictions on the number or size of assistant-related objects. Limits below are scoped per organization except for Assistants per project, which is scoped per project.
*1,000,000 input tokens/month to explore Marketplace apps until June 30, 2026.
Additionally, the following limits apply to multimodal PDFs (currently in public preview):
Multimodal PDF processing uses the same ingestion unit as standard uploads; it is billed at about twice the standard per-unit rate (see Pricing and limits). Object and rate limits for assistants also applyāsee #limits and #rate-limits.
Rate limits
Rate limits help protect your applications from misuse and maintain the health of our shared infrastructure. These limits are designed to support typical production workloads while ensuring reliable performance for all users. Most rate limits can be adjusted upon request. If you need higher limits to scale your application, contact Support with details about your use case. Requests that exceed a rate limit fail and return a429 - TOO_MANY_REQUESTS status.