Skip to main content
TruLens is an open-source library for evaluating and tracking large language model (LLM) applications. It scores each request with feedback functions, such as answer relevance and context relevance, and records latency and cost, so you can compare app configurations as you iterate.

Setup guide

This guide builds a retrieval-augmented generation (RAG) app on Pinecone Database with LangChain, evaluates it with TruLens, and compares configurations.
1

Set up the environment

Install the libraries, and set your OpenAI API key as the OPENAI_API_KEY environment variable. LangChain and the TruLens OpenAI provider both read it.
Shell
2

Create the index

Load a pre-embedded dataset from pinecone-datasets, so you can skip the embedding step:
Python
Create an index and upsert the documents. The dataset was embedded with text-embedding-ada-002, so the index has 1536 dimensions. The distance metric is the first configuration choice you compare later.
Python
3

Build the RAG chain

Create a LangChain vector store on the index, and build a chain that retrieves documents and passes them to the LLM with the question. Queries must use the same embedding model as the dataset.
Python
4

Define feedback functions

Feedback functions score each request. This guide uses two:
  • Context relevance is the average relevance (0 to 1) of each context chunk the retriever returns.
  • Answer relevance is the relevance (0 to 1) of the final answer to the question.
Python
Selectors tell TruLens which parts of a request to score. select_record_input() and select_record_output() are the app’s question and final answer. select_context(collect_list=False) is each context chunk the retriever returns, scored separately, and agg=np.mean averages those scores.
5

Record queries

Wrap the chain with TruLens and run queries inside the recorder’s context. TruLens computes feedback in the background, so retrieve_feedback_results waits for the scores and returns one row per query.
Python
Compare the scores, latency, and cost of each app version with the leaderboard, or explore individual records in the TruLens dashboard:
Python
6

Compare configurations

To compare configurations, change one component, wrap the new chain with a new app_version, and run the same queries. Then check the leaderboard again.The distance metric is set when you create the index. To try euclidean or dotproduct, create and populate a second index, and point a new vector store and chain at it. OpenAI embeddings are normalized to length 1, so all three metrics return the same ranking, and any difference shows up in latency rather than quality.
Python
To try a different model or a different amount of context (top k), swap the LLM or set k on the retriever:
Python
After each change, rebuild the chain, wrap it with a new version, and record the same queries. Use a new variable for each version, so earlier recorders keep running until their feedback finishes.
Python
7

Clean up

When you’re finished with the index, delete both indexes.
Python