Setup guide
This guide shows how to integrate Pinecone and the Haystack library for question answering.Install Haystack
Install the latest version of Haystack with all dependencies required for thePineconeDocumentStore.
Python
Initialize the PineconeDocumentStore
Initialize aPineconeDocumentStore by providing an API key and environment name. Create an account to get your free API key.
Python
Prepare data
Before you add data to the document store, you must download the data and convert it into the Document format that Haystack uses. This guide uses the SQuAD dataset available from Hugging Face Datasets.Python
Python
Then convert these records into the Document format.
Python
Document format contains two fields: content for the text content or paragraphs, and meta for any additional information you can later use to apply metadata filtering in your search.
Upsert the documents to Pinecone.
Python
Initialize retriever
The next step is to create embeddings from these documents. This guide uses Haystack’sEmbeddingRetriever with a SentenceTransformer model (multi-qa-MiniLM-L6-cos-v1), which is designed for question answering.
Python
PineconeDocumentStore.update_embeddings method with the retriever provided as an argument. GPU acceleration can greatly reduce the time required for this step.
Python
Inspect documents and embeddings
You can get documents by their ID with thePineconeDocumentStore.get_documents_by_id method.
Python
d.content and the document embedding with d.embedding.
Initialize an extractive QA pipeline
AnExtractiveQAPipeline contains three key components by default:
- a document store (
PineconeDocumentStore) - a retriever model
- a reader model
deepset/electra-base-squad2 model from the Hugging Face model hub as the reader model.
Python
ExtractiveQAPipeline.
Python
Ask questions
Use your QA pipeline to start querying withpipe.run.
Python
Python
Python
top_k parameter.
Python