Setup guide
In this guide, you’ll create embeddings based on the sentence-transformers/all-MiniLM-L6-v2 model from Hugging Face, but the approach demonstrated here should work with any other model and dataset.Before you begin
Ensure you have the following:Install the Spark-Pinecone connector
- Databricks platform
- Databricks on AWS
- Databricks on GCP / Azure
- Install the Spark-Pinecone connector as a library.
- Configure the library as follows:
- Select File path/S3 as the Library Source.
-
Enter the S3 URI for the Pinecone assembly JAR file:
Databricks platform users must use the Pinecone assembly jar listed above to ensure that the proper dependecies are installed.
- Click Install.
Load the dataset into partitions
This guide uses a collection of news articles from the Hugging Face Datasets library as the example dataset. To load it, follow these steps:
- Create a new notebook attached to your cluster.
-
Install dependencies:
Shell
-
Load the dataset:
Python
-
Convert the dataset from the Hugging Face format and repartition it:
Once the repartition is complete, you get back a DataFrame, which is a distributed collection of the data organized into named columns. It’s conceptually equivalent to a table in a relational database or a dataframe in R/Python, but with richer optimizations under the hood. Each partition in the DataFrame has an equal amount of the original data.Python
-
The dataset doesn’t have identifiers associated with each document, so add them:
As its name suggests,Python
withColumnadds a column to the dataframe, containing a simple increasing identifier that you cast to a string.
Create embeddings
Generate an embedding for each document, and then convert the results to the schema Pinecone expects:
-
Create a user-defined function (UDF) to create the embeddings, using the AutoTokenizer and AutoModel classes from the Hugging Face transformers library:
Python
-
Apply the UDF to the data:
A dataframe in Spark is a higher-level abstraction built on top of a more fundamental building block called a resilient distributed dataset (RDD). Here, you use thePython
mapPartitionsfunction, which provides finer control over the execution of the UDF by explicitly applying it to each partition of the RDD. -
Convert the resulting RDD back into a dataframe with the schema required by Pinecone:
Python
Store the embeddings
Write the embeddings to a Pinecone index with the Spark-Pinecone connector, and then query the index:
-
Initialize the connection to Pinecone:
Python
-
Create an index for your embeddings. The
dimensionmust match the model’s output size, which is 384 for all-MiniLM-L6-v2:Python -
Use the Spark-Pinecone connector to save the embeddings to your index:
Python
-
Perform a similarity search using the embeddings you loaded into Pinecone by providing a set of vector values or a vector ID. The query endpoint returns the IDs of the most similar records in the index, along with their similarity scores:
PythonIf you want to make a query with a text string (e.g.,
"Summarize this article"), use thesearchendpoint via integrated inference.