Context Relevance Score
The Context Relevance Score is a metric (0-1) that measures the semantic alignment between a webpage's content and its stated topic or the likely intent of a user query.
For RAG systems, retrieving highly relevant context is the most critical step; irrelevant or contextually drifted information leads directly to poor or incorrect generated answers. This score uses vector embeddings to move beyond simple keyword matching and quantify the conceptual coherence of a page. A high score indicates that the content is tightly focused on its purported subject, making it a reliable and efficient source for retrieval.
Calculation Methodology
The score is calculated by comparing the semantic meaning of the page's topic signifiers with the semantic meaning of its body content.
Topic Vector Generation
- Scrape the page's primary topic indicators: the <title> tag and the <h1> tag.
- Concatenate these two strings to form a "topic document."
- Using a sentence-transformer model (e.g., from the sentence-transformers library), generate a dense vector embedding for this topic document. This vector, Vtopic, numerically represents the intended subject of the page.
Content Vector Generation
- Scrape the text from the main content area of the page (e.g., within <main> or <article> tags).
- Generate a dense vector embedding for the entire body of content. This vector, Vcontent, represents the actual subject matter discussed.
Primary Relevance Calculation
Calculate the cosine similarity between the two vectors, Vtopic and Vcontent. Cosine similarity measures the cosine of the angle between two vectors, resulting in a score between -1 and 1 (though typically 0 to 1 for this use case). The formula is: Similarity = ||Vtopic|| * ||Vcontent|| / (Vtopic ⋅ Vcontent). A score close to 1 indicates a very high degree of semantic overlap and strong context relevance.
Query-Intent Simulation (Optional Enhancement)
To further validate relevance, use an LLM to generate a set of 5-10 plausible user queries that the page aims to answer. Generate embeddings for each of these simulated queries and calculate their average cosine similarity to the content vector. A high score here indicates the page effectively addresses likely user intents.
Calculating The Context Relevance Score
The calculation of the Context Relevance Score is a multi-step process that involves analyzing the content and the author's background. Here is a simplified pseudo-code representation of how this score is calculated.
Pseudo-code for Context Relevance Score Calculation
BEGIN FETCH and PARSE the webpage content. EXTRACT topic_text from the <title> and <h1> tags. EXTRACT content_text from the main body. GENERATE a vector embedding for topic_text -> V_topic. GENERATE a vector embedding for content_text -> V_content. CALCULATE the cosine similarity between V_topic and V_content to get primary_relevance_score. // Optional Enhancement USE LLM to GENERATE a list of simulated_user_queries based on topic_text. GENERATE vector embeddings for each query. CALCULATE the average cosine similarity between V_content and the query vectors to get query_relevance_score. COMBINE primary_relevance_score and query_relevance_score using weights to get final_score. RETURN final_score. END
Conclusion
By focusing on the Context Relevance Score, you are ensuring that your content is not just visible, but is also understood and valued by AI systems. This is a critical step in building a strong foundation for your GEO strategy.
Key Takeaways
- The Context Relevance Score measures the semantic alignment of your content with user intent.
- A high score indicates that your content is a good source for RAG systems.
- The calculation involves comparing vector embeddings of the topic and content.
- You can enhance the score by simulating user queries and checking for alignment.