Skip to main content
Hugging Face Inference Endpoints offers a secure production solution for deploying any Hugging Face Transformers, Sentence-Transformers, and Diffusion models from the Hub on dedicated, autoscaling infrastructure managed by Hugging Face. With Pinecone, you can use Hugging Face to generate and index vector embeddings.

Setup guide

Hugging Face Inference Endpoints provides access to model inference. In this guide, you create an endpoint that generates vector embeddings, index the embeddings in Pinecone, and search them.
1

Create an endpoint

In Hugging Face Inference Endpoints, create an endpoint for a sentence-embedding model from the Hub, such as sentence-transformers/all-mpnet-base-v2, which outputs 768-dimensional embeddings like the examples in this guide. A CPU instance costs less, and a GPU instance embeds faster. Endpoints are billed while they run, so delete yours when you finish.When the endpoint’s status is Running, copy its endpoint URL from the endpoint’s overview page:
Python
You also need a Hugging Face access token that can call your endpoints. Create a fine-grained token in your Hugging Face account settings under Access Tokens:
Python
2

Create embeddings

Define a helper that sends text to the endpoint and returns the embeddings. Depending on the serving engine, an endpoint returns either {"embeddings": [...]} or a bare list of embeddings, so the helper handles both.
Python
Response
Check the dimensionality of the embeddings. You need it when you create the index.
Python
Response
You need more than two items to search through, so download a larger dataset with Hugging Face Datasets.
Python
Response
SNLI contains 550K sentence pairs, and many of them include duplicate items, so take just one set of these (the hypothesis column) and deduplicate it.
Python
Response
To keep the example quick to run, reduce the set to 50K sentences. If you have time, you can keep the full 480K.
Python
3

Create an index

With your endpoint and dataset ready, all that’s missing is a Pinecone index. First, initialize your connection to Pinecone, which requires an API key.
Python
Now create a new index called 'hf-endpoints'. You can use any name. The dimension must match your endpoint model’s output dimensionality (which you found in dim earlier), and the metric must match the model (cosine typically works, but not for all models).
Python
4

Store the embeddings

With the endpoint, dataset, and Pinecone index ready, create embeddings for the dataset and index them in Pinecone.
Python
Response
6

Clean up

Shut down the endpoint by navigating to the Inference Endpoints Overview page and selecting Delete endpoint. Delete the Pinecone index with:
Python
Once the index is deleted, you can’t use it again.