Setup guide
This guide builds a retrieval-augmented generation (RAG) app on Pinecone Database with LangChain, evaluates it with TruLens, and compares configurations.Set up the environment
Install the libraries, and set your OpenAI API key as the
OPENAI_API_KEY environment variable. LangChain and the TruLens OpenAI provider both read it.Shell
Create the index
Load a pre-embedded dataset from Create an index and upsert the documents. The dataset was embedded with
pinecone-datasets, so you can skip the embedding step:Python
text-embedding-ada-002, so the index has 1536 dimensions. The distance metric is the first configuration choice you compare later.Python
Build the RAG chain
Create a LangChain vector store on the index, and build a chain that retrieves documents and passes them to the LLM with the question. Queries must use the same embedding model as the dataset.
Python
Define feedback functions
Feedback functions score each request. This guide uses two:Selectors tell TruLens which parts of a request to score.
- Context relevance is the average relevance (0 to 1) of each context chunk the retriever returns.
- Answer relevance is the relevance (0 to 1) of the final answer to the question.
Python
select_record_input() and select_record_output() are the app’s question and final answer. select_context(collect_list=False) is each context chunk the retriever returns, scored separately, and agg=np.mean averages those scores.Record queries
Wrap the chain with TruLens and run queries inside the recorder’s context. TruLens computes feedback in the background, so Compare the scores, latency, and cost of each app version with the leaderboard, or explore individual records in the TruLens dashboard:
retrieve_feedback_results waits for the scores and returns one row per query.Python
Python
Compare configurations
To compare configurations, change one component, wrap the new chain with a new To try a different model or a different amount of context (top k), swap the LLM or set After each change, rebuild the chain, wrap it with a new version, and record the same queries. Use a new variable for each version, so earlier recorders keep running until their feedback finishes.
app_version, and run the same queries. Then check the leaderboard again.The distance metric is set when you create the index. To try euclidean or dotproduct, create and populate a second index, and point a new vector store and chain at it. OpenAI embeddings are normalized to length 1, so all three metrics return the same ranking, and any difference shows up in latency rather than quality.Python
k on the retriever:Python
Python