# Build an agentic knowledge graph that cites every claim

A baseline RAG pipeline scores 45.6 percent exact match on HotpotQA and 10.8 percent joint score, because the second document a question depends on never reaches the context window (Yang and colleagues, 2018).

**In a nutshell.** An agentic knowledge graph links entities to the text chunks that mention them, so retrieval follows relational edges instead of guessing which passage looks similar to the query. On the HotpotQA distractor setting, the original baseline reaches 59.0 F1 on the answer but only 40.2 on the joint answer-and-evidence score, and the fullwiki setting drops to 16.1 joint F1 (Yang and colleagues, 2018).

Raising top-k does not close that gap, because it widens the prompt instead of adding the missing hop., and in this guide builds the missing structure in six phases: streaming ingestion, hybrid retrieval, evidence-linked extraction, a graph stored inside PostgreSQL with Apache AGE, a bounded agent with four read-only tools, and a verification gate that abstains rather than guesses.

  SOVEREIGN EUROPEAN INFERENCE 

###  Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention 

 Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.

 [ Start 30-day trial → ](https://regolo.ai/pricing/?utm_source=blog&utm_medium=bento_cta&utm_campaign=deepseek-flash-mid)  Free credits included · Live in 60s  

  ✓ 100% EU Green Datacenters   ✓ Certified ZDR  

 

 

## Why standard vector search fails: the multi-hop retrieval bottleneck

Standard retrieval-augmented generation assumes that semantic similarity to a user query will pull all required context into the prompt, when an inquiry requires facts from two different documents, that assumption breaks down.

![](https://regolo.ai/wp-content/uploads/2026/10/Traditional-Vector-RAG-vs.-Agentic-GraphRAG-927x1024.png)The original HotpotQA baseline quantifies the damage. In the distractor setting it reaches 45.6 answer exact match and 59.0 F1, but joint exact match — the answer *and* the full annotated evidence set correct together — collapses to 10.8, and joint F1 to 40.2. On the fullwiki setting, where supporting passages must be retrieved from all of Wikipedia rather than handed to the model, joint F1 falls to 16.1 (Yang and colleagues, 2018).

That gap is the whole problem – a model can be handed all ten distractor paragraphs and still fail, because the second required passage is present in the context but not selected as supporting evidence. When the passage is missing entirely, no amount of reasoning recovers it.

Dense vector search locates isolated passages with high semantic similarity, but it cannot follow relational connections between entities. It evaluates each document as an isolated text block, so it has no way to represent that one passage is a prerequisite for another.

### The fundamental assumption of vector search

Vector search models convert text into high-dimensional numerical vectors. A retrieval engine computes cosine similarity between the query vector and precomputed document vectors:

```
similarity(query, document) = (query · document) / (||query|| * ||document||)Code language: JavaScript (javascript)
```

This mathematical approach works well when user queries share topical vocabulary or semantic meaning with a single target passage. The model ranks documents by topical closeness.

However, cosine similarity cannot evaluate logical dependencies across documents. It is a measure of resemblance between two text fragments in isolation, and a document about corporate hotel headquarters in Delhi bears little resemblance to a query about the Oberoi family.

### The concrete failure case: HotpotQA

Consider this multi-hop benchmark question from HotpotQA (Yang and colleagues, 2018):

> The Oberoi family is part of a hotel company that has a head office in what city?

Answering this question requires two distinct pieces of evidence from separate Wikipedia passages:

- document a ("Oberoi family"): states that the family controls The Oberoi Group.
- document b ("The Oberoi Group"): states that the corporate headquarters is located in Delhi.

The language model needs both passages to produce the correct answer, Delhi.

HotpotQA evaluates two types of multi-document questions: a **bridge question** connects a chain of facts across documents, such as finding the city where a company founded by the Oberoi family has its headquarters – and a **comparison question** compares properties across separate entities, such as deciding whether Arthur or George lived longer.

### What happens during vector retrieval

When the user query enters the retrieval system, the embedding model converts the text into a single vector; the phrase "Oberoi family" dominates the query representation.

The retrieval engine matches the query against the vector index:

1. Document a contains the exact phrase "Oberoi family": the cosine similarity score is high, document a enters the candidate context.
2. Document b does not mention the phrase "Oberoi family": It describes corporate hotel management and lists cities. As a result, document b produces a low similarity score against the query vector.
3. The retrieval engine truncates results at rank k (typically 5 or 10 passages), document b falls outside the top-k threshold.

Because document b is omitted, the prompt context contains only document a. The language model cannot locate the corporate headquarters. The model either hallucinates an incorrect city or states that it cannot answer.

### How expensive the miss actually is

When no coverage gap occurs — every required hop reached the context — average accuracy is 77.9 percent, instead when a coverage gap exists, accuracy collapses to 49.2 percent, a penalty of 28.7 percentage points.

Once the retriever fails to surface a single required hop, most models cannot reason their way to the correct answer.

---

## Why common workarounds fail

Engineers typically attempt two workarounds to resolve this retrieval failure.

### Workaround 1: increasing the retrieval limit (top-k)

Engineers often expand k from 5 to 50 chunks to capture missing documents. This strategy creates three operational bottlenecks:

- **context window bloat**: passing 50 passages into the prompt increases token consumption and API expenses tenfold.
- **needle in a haystack degradation**: language models struggle to extract relevant facts when surrounded by irrelevant distractor text.
- **high latency**: embedding lookups and larger prompt payloads increase overall response time.

Even with k set to 50, document b can remain unretrieved if the database contains dozens of generic hotel articles with higher similarity scores. Top-k raises the ceiling, and It does not change the ranking function, and the ranking function is what failed.

### Workaround 2: naive iterative re-prompting

Some systems prompt the model to generate secondary search queries automatically without an explicit entity graph, this approach suffers from severe limitations:

- **blind query formulation**: the model must guess search terms without structural constraints.
- **latency accumulation**: each iterative model call adds network round-trips and token expenses.
- **termination ambiguity**: the system lacks deterministic criteria to decide when retrieved context is complete.

## The structural solution: agentic graphrag

The fix is to stop asking which passage resembles the query, and start recording which passage asserts something *about* which entity. Once edges point from entities to the chunk that supports them, multi-hop evidence becomes reachable by traversal rather than by luck.

By linking graph edges directly to text chunk identifiers, the system discovers multi-hop evidence without relying on accidental semantic similarity.

Published GraphRAG results deserve a careful reading, because the headline numbers attached to graph retrieval are frequently softer than they look. What Edge and colleagues (2024) actually measure is comparative win rates against a baseline RAG system on query-focused summarization, not a single accuracy delta. Intermediate-level community summaries in the Podcast dataset and low-level community summaries in the News dataset achieved comprehensiveness win rates of 57 percent (p &lt; .001) and 64 percent (p &lt; .001); diversity win rates were 57 percent (p = .036) and 60 percent (p &lt; .001). Root-level GraphRAG retains advantages over vector RAG at a 72 percent comprehensiveness win rate and 62 percent diversity win rate.

![](https://regolo.ai/wp-content/uploads/2026/10/Screenshot-2026-10-09-alle-09.59.27-1024x281.png)![](https://regolo.ai/wp-content/uploads/2026/10/Screenshot-2026-10-09-alle-09.59.36-1024x307.png)IMAGE: how to map the entire data flow, from raw documents to the verified response.

The larger and more defensible effect in the same paper is cost. For low-level community summaries, GraphRAG required 26 to 33 percent fewer context tokens; for root-level community summaries, over 97 percent fewer tokens than source text summarization. A 97 percent token reduction that holds a 72 percent comprehensiveness win rate is a trade most production teams should take without hesitating.

Treat any single headline accuracy percentage attached to graph retrieval with suspicion, including figures in vendor case studies. A 35 percent improvement claim from a vendor partner of a cloud provider is a marketing number, not a benchmark result.

## Choose evidence-aligned datasets

A large text collection does not prove that graph traversal improves answer quality. You need questions that require multiple sources, explicit annotations for supporting facts, and a corpus large enough to measure ingestion latency.

Our primary benchmark dataset is HotpotQA (Yang and colleagues, 2018):

| Dataset | Role in this guide | Provided data |
|---|---|---|
| `hotpotqa/hotpot_qa` | Initial testing and answer evaluation | Multi-hop questions, answers, and sentence annotations |
| `BeIR/hotpotqa` | Corpus scaling and retrieval tests | 5,233,329 corpus passages and 97,852 queries |
| `BeIR/hotpotqa-qrels` | Retrieval quality scoring | Query-to-document relevance assessments |
| `framolfese/2WikiMultihopQA` | Supplemental graph tests | Multi-hop questions with structured relationship annotations |

### Separate operational data from evaluation labels

`bench/prepare.py` generates two artifacts:

- `corpus.parquet`: source identifiers, titles, text passages, and document provenance.
- `evaluation.jsonl`: questions, target answers, supporting documents, and supporting sentence identifiers.

Only the corpus enters the ingestion pipeline. Target answers and supporting sentence annotations remain in the evaluation harness. Those labels never enter extraction prompts, graph construction logic, or agent tools.

`bench/align_hotpot.py` checks mappings between HotpotQA questions, BEIR queries, relevance assessments, and source passages. The script matches question identifiers explicitly instead of assuming that identical split names contain identical records.

The alignment report records matched questions, missing support passages, ambiguous titles, and exact sentence matches. If sentence alignment fails, we can report document retrieval metrics, but we will not report sentence-level supporting fact scores.

### Pin revisions and preserve attribution

`datasets.lock.json` stores exact Hugging Face commit hashes and source metadata. You must not use an unpinned main branch for a published benchmark.

HotpotQA and its processed Wikipedia corpus use the CC BY-SA 4.0 license. Dataset attribution and legal obligations remain distinct from the software license of the application. You must download the full corpus through preparation scripts rather than storing it in the code repository.

---

  SOVEREIGN EUROPEAN INFERENCE 

###  Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention 

 Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.

 [ Start 30-day trial → ](https://regolo.ai/pricing/?utm_source=blog&utm_medium=bento_cta&utm_campaign=deepseek-flash-mid)  Free credits included · Live in 60s  

  ✓ 100% EU Green Datacenters   ✓ Certified ZDR  

 

 

## Getting started

You need Python 3.11 or later, a [Regolo API key](/signup), and an open-source container runtime such as Docker engine.

The `compose.yaml` configuration in the repository launches PostgreSQL with Apache AGE and Qdrant with persistent disk volumes. Database container images use tested version tags and image digests instead of the latest tag.

### Configure the environment

Copy `.env.example` to `.env` and define these environment variables:

```
REGOLO_API_BASE=https://api.regolo.ai/v1
REGOLO_API_KEY=replace_me

REGOLO_CHAT_MODEL=gpt-oss-120b
REGOLO_EMBED_MODEL=Qwen3-Embedding-8BCode language: Bash (bash)
```

Our embedding requests target `/v1/embeddings`. Model discovery uses the `/models` and `/model_group/info` endpoints on the api domain. You must not assume that every administrative route uses the `/v1` prefix.

### Run preflight checks

`app/preflight.py` validates model access, required tool capabilities, a live chat request, a live embedding request, vector dimensions, and database connections.

The model catalog check runs through this request logic:

```
response = http.get(
    "https://api.regolo.ai/model_group/info",
    headers={"Authorization": f"Bearer {settings.regolo_api_key}"},
)
response.raise_for_status()

catalog = {item["model_group"]: item for item in response.json()["data"]}
chat = catalog[settings.chat_model]

if chat["mode"] != "chat" or not chat.get("supports_function_calling"):
    raise RuntimeError("Configured model does not meet agent requirements")Code language: Python (python)
```

The preflight script also validates structured function calls. You must not infer tool support from a model family name alone.

### Run the demo workflow

When you initialize the development repository, run this sequence of commands:

```
cp .env.example .env
# Edit .env and insert your API key

python -m venv .venv
source .venv/bin/activate
pip install -e ".[bench]"

docker compose up -d
python -m app.preflight
python -m app.db.migrate

python -m bench.prepare --profile demo --questions 100 --seed 42
python -m app.ingest --profile demo
python -m app.ask --query-id <QUESTION_ID_FROM_THE_DEMO_MANIFEST>Code language: Bash (bash)
```

The preparation script generates a manifest that contains selected question identifiers. `app.ask` loads a target question from that manifest without exposing evaluation labels to the model.

---

## Build the assistant

Each phase contains a small set of files, an observable output, and an explicit verification step before subsequent steps.

### Phase 1: read and normalize documents

Files: `app/datasets/hotpot.py`, `app/normalization.py`, and `app/ingest.py`.

For large datasets, you must stream records to avoid excessive memory consumption:

```
from datasets import load_dataset

corpus = load_dataset(
    "BeIR/hotpotqa",
    "corpus",
    split="corpus",
    revision=dataset_revision,
    streaming=True,
)

for record in corpus:
    yield {
        "source_id": str(record["_id"]),
        "title": record["title"],
        "text": record["text"],
    }Code language: Python (python)
```

The dataset provides the corpus configuration with `_id`, `title`, and `text` fields. In the demo adapter, the pipeline preserves sentence boundaries and sentence indices from HotpotQA.

Short passages remain intact. Longer passages pass through `app/chunking.py`, which preserves sentence offsets so that child blocks map back to source documents.

Documents receive stable source identifiers, versions receive content hashes, and chunks receive identifiers based on document, version, and text span. Identical text does not justify deleting BEIR passage identifiers. You can cache embeddings by content hash while retaining original record identifiers for relevance scoring.

Make sure that you inspect counts for accepted records, empty inputs, rejected items, and retained chunks.

### Phase 2: build hybrid retrieval

Files: `app/regolo_client.py`, `app/vector_store.py`, `app/retrieval.py`, and `app/db/migrations/002_search.sql`.

PostgreSQL locates exact names and keywords, while Qdrant locates semantically related passages. The application generates all vector embeddings through Regolo:

```
response = await http.post(
    f"{settings.regolo_api_base}/embeddings",
    headers=auth_headers,
    json={"model": settings.embed_model, "input": texts},
)
response.raise_for_status()

items = sorted(response.json()["data"], key=lambda item: item["index"])
vectors = [item["embedding"] for item in items]Code language: Python (python)
```

Batch limits depend on the selected model. You must make sure that item counts and vector dimensions match the preflight values. You must not truncate passage text silently.

For lexical search, create a generalized inverted index (GIN, an index structure for fast text lookups) in PostgreSQL:

```
CREATE INDEX chunks_search_idx
ON app.chunks
USING GIN (
    to_tsvector('english', coalesce(title, '') || ' ' || text)
);Code language: SQL (Structured Query Language) (sql)
```

The retrieval query uses the same text configuration and search expression. This mechanism is PostgreSQL full-text ranking rather than BM25.

Qdrant points use deterministic identifiers and payload attributes such as `document_id`, `version`, and `chunk_id`. The application fuses lexical and semantic rankings with reciprocal rank fusion (RRF, an algorithm that combines rank scores across search systems):

```
for ranking in (lexical_results, semantic_results):
    for rank, hit in enumerate(ranking, start=1):
        scores[hit.chunk_id] += 1.0 / (rrf_constant + rank)Code language: Python (python)
```

The fusion constant, candidate limits, and evidence quotas are configuration parameters chosen on development data. Make sure that you inspect individual rankings and fused scores on sample questions.

### Phase 3: extract supported relationships

Files: `app/extraction.py`, `app/schemas.py`, and `app/entity_resolution.py`.

The extraction prompt takes a passage title and numbered sentences. It enforces strict constraints:

```
Extract only relationships explicitly supported by this passage.
Return subject, predicate, object, sentence_ids, and evidence_quote.
Do not use outside knowledge.
Do not invent relationships between documents.Code language: Bash (bash)
```

A pydantic schema validates the model response:

```
class extracted_relation(BaseModel):
    subject: str
    predicate: str
    object: str
    sentence_ids: list[int]
    evidence_quote: strCode language: Python (python)
```

The schema rejects unexpected fields and enforces length constraints on predicates and identifiers. The extraction code verifies that cited sentence identifiers exist and that quoted text appears in the source passage.

An exact quote is necessary but not sufficient, because a model can cite a real sentence while misunderstanding its meaning. You must retain evidence for downstream verification and review a sample of extracted edges manually.

Entity resolution must remain conservative. Matching names do not establish identical real-world identities. The resolver tracks canonical identifiers, aliases, and unresolved ambiguities without merging records based solely on text labels.

Make sure that you track accepted relationships, rejected extractions, and ambiguous matches.

### Phase 4: store a persistent evidence-linked graph

Files: `app/graph_store.py` and `app/db/migrations/003_graph.sql`.

PostgreSQL stores authoritative facts with provenance links, while Apache AGE stores a navigable graph representation:

```
(Entity)-[:FACT {
    fact_id,
    predicate,
    source_chunk_id
}]->(Entity)Code language: Bash (bash)
```

Attributes such as dates and numerical measurements can use value nodes. This structure supports comparison questions without forcing values into entity nodes.

![](https://regolo.ai/wp-content/uploads/2026/10/store-a-persistent-evidence-linked-graph-1024x981-1.png)A single relationship can reference multiple supporting passages. Contradictory claims preserve separate provenance records instead of overwriting existing edges. Stable fact identifiers support retry operations and database reconciliation.

Initialize Apache AGE with this migration script:

```
CREATE EXTENSION IF NOT EXISTS age;
LOAD 'age';
SET search_path = ag_catalog, "$user", public;

SELECT create_graph('knowledge');Code language: SQL (Structured Query Language) (sql)
```

The migration checks whether the graph exists before execution. Connection pools must execute `LOAD` and configure `search_path` for every active connection. Writes must commit before worker processes can read updated graph data.

To resolve the Oberoi question, the application executes this openCypher query inside PostgreSQL:

```
SELECT * FROM cypher('knowledge', $$
    MATCH (e1:Entity {name: 'Oberoi family'})-[r:FACT]->(e2:Entity)
    RETURN e1.name, r.predicate, e2.name, r.source_chunk_id
$$) as (entity_1 agtype, predicate agtype, entity_2 agtype, chunk_id agtype);Code language: SQL (Structured Query Language) (sql)
```

When this query returns `The Oberoi Group` and its `source_chunk_id`, the system retrieves the original text chunk from PostgreSQL with a standard query:

```
SELECT chunk_id, title, text
FROM app.chunks
WHERE chunk_id = 'chunk_oberoi_group_01';Code language: SQL (Structured Query Language) (sql)
```

The model never generates raw SQL or Cypher strings. The graph adapter uses bounded query templates and retrieves immediate neighbors with at most two expansion levels in the initial configuration.

Make sure that every edge links to valid evidence, retries do not duplicate rows, and the graph remains intact after a database restart.

### Phase 5: add bounded agent behavior

Files: `app/tools.py`, `app/agent.py`, and `app/prompts/agent.txt`.

The agent receives four read-only tools:

```
tools = {
    "search_documents": search_documents,
    "lookup_entity": lookup_entity,
    "expand_graph": expand_graph,
    "get_evidence": get_evidence,
}Code language: JavaScript (javascript)
```

The controller validates tool arguments before execution. The model selects subsequent actions dynamically rather than following a fixed pipeline.

In the Oberoi inquiry, the model searches for the family, identifies the company entity, traverses the corporate link, and retrieves the headquarters passage. If the graph lacks a connection, the agent can fall back to text search using the company name.

![](https://regolo.ai/wp-content/uploads/2026/10/Screenshot-2026-10-09-alle-10.01.04-1024x527.png)This fallback path is essential. When an agent succeeds through keyword search, the evaluation harness must not credit that success to graph traversal. The system records which tools and evidence chunks contributed to the final answer.

The reference configuration allows six tool calls, two graph expansion steps, and a fixed evidence budget. Execution deadlines and token limits constrain total resource usage.

Make sure that you record tool names, call arguments, retrieved evidence identifiers, fallback events, and budget exhaustion states.

### Phase 6: make sure that evidence supports claims before answering

![](https://regolo.ai/wp-content/uploads/2026/10/Screenshot-2026-10-09-alle-10.01.29-347x1024.png)

Files: `app/answering.py` and `app/verification.py`.

The response structure separates the short answer from explanatory claims:

```
{
  "answer": "Delhi",
  "claims": [
    {
      "text": "The company has its head office in Delhi.",
      "evidence_ids": ["chunk_b:sentence_0"]
    }
  ],
  "status": "answered"
}Code language: JavaScript (javascript)
```

Deterministic validation checks evidence identifiers, chunk versions, access controls, and quote offsets. A separate Regolo inference request evaluates whether cited passages support each claim.

The publication rule enforces strict verification:

```
if not citations_valid or not claims_supported:
    return abstain(reason="insufficient_or_unsupported_evidence")Code language: Python (python)
```

The assistant can return a verified answer, a verified partial answer, or an abstention. The system retains unverified claims in the audit log instead of deleting them silently.

Evaluating claims with the generator model does not provide independent verification. Verification is a pipeline component under test rather than an authoritative benchmark judge.

Retrieved document text represents untrusted data. Input passages must not activate tools, alter permissions, access benchmark labels, or override agent operating instructions.

Make sure that you test missing evidence, invalid citations, contradictory source texts, and adversarial instructions embedded in documents.

---

## Evaluate and scale

### Phase 7: measure the system and expose its API

Files: `bench/evaluate_retrieval.py`, `bench/evaluate_answers.py`, `bench/load_test.js`, `bench/report.py`, and `app/api.py`.

Before you measure execution speed, you must isolate the factual contribution of the knowledge graph:

| Configuration | Operational objective |
|---|---|
| Hybrid retrieval and answer | Baseline single-pass pipeline |
| Hybrid retrieval, fixed graph expansion, and answer | Measure fixed graph expansion contribution |
| Search-only agent without graph tools | Measure iterative search contribution |
| Complete agent with graph tools and verification | Measure full system capability |

Without this baseline, additional search calls can be mistaken for benefits of graph traversal, you must keep the model, dataset, and context quotas consistent across all configurations. Prompt tuning happens on development splits, while final benchmark questions remain separate.

### Measure retrieval completeness

Recall at k can appear high even when a question lacks its second required source document. Alongside Recall@k and ndcg@10 (normalized discounted cumulative gain at rank 10, a metric for ranked search relevance), **you must measure whether the retriever recovered all annotated support passages:**

```
def all_support_retrieved(
    retrieved_ids: list[str], gold_ids: list[str], k: int
) -> float:
    gold = set(gold_ids)
    if not gold:
        raise ValueError("Missing relevance annotations")
    return float(gold.issubset(set(retrieved_ids[:k])))Code language: Python (python)
```

Chunk identifiers map back to original corpus passage identifiers before evaluation. This metric scores coverage of labeled evidence rather than claiming that all unannotated passages are irrelevant.

### Measure answers and evidence

We calculate exact match (EM) and F1 scores for answers, supporting fact EM and F1 where sentence alignment matches, and joint accuracy metrics. The HotpotQA project provides its official evaluation script, we follow its scoring format while labeling reduced-corpus experiments as a distinct evaluation protocol.

### You must report bridge questions and comparison questions separately, the answer field contains the short entity name rather than an explanatory paragraph, so the scoring script compares exact terms.

A closed-book baseline without documents reveals questions that the model remembers from pre-training, but it does not rule out contamination. Evidence recovery and evaluation on proprietary documents remain necessary.

To evaluate abstention behavior, construct a controlled test set by removing necessary evidence passages. This test provides a resilience benchmark rather than an official HotpotQA metric.

### Increase corpus size systematically

You must test 10,000, 100,000, and 1,000,000 passages before indexing the full corpus of 5.23 million records; reduced evaluation runs must retain all support documents for selected queries.

Every run report must include passage totals, chunk counts, vector dimensions, entity counts, relationship counts, and the fraction of the corpus represented in the graph. A complete vector index paired with a partial graph does not constitute a complete knowledge graph.

  SOVEREIGN EUROPEAN INFERENCE 

###  Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention 

 Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.

 [ Start 30-day trial → ](https://regolo.ai/pricing/?utm_source=blog&utm_medium=bento_cta&utm_campaign=deepseek-flash-mid)  Free credits included · Live in 60s  

  ✓ 100% EU Green Datacenters   ✓ Certified ZDR  

 

 

---

## Frequently asked questions

### What is an agentic knowledge graph?

An agentic knowledge graph connects structured relational entities to original text documents while an autonomous controller executes iterative search and graph traversal tools to collect multi-hop evidence.

### How does graphrag differ from standard vector rag?

Vector rag retrieves isolated passages through semantic text similarity, whereas graphrag discovers connections between distinct entities mentioned across separate documents.

### Why use apache age with postgresql instead of a standalone graph database?

Apache AGE keeps relational documents, transactional state, and graph topologies within PostgreSQL, which avoids maintaining separate operational databases for transactional metadata.

### How does the system prevent hallucinated graph connections?

The extraction pipeline requires sentence identifiers and exact evidence quotes for every relation, while the controller verifies citations before publishing an answer.

### When must you use hybrid retrieval over vector search alone?

You must use hybrid retrieval when user queries include specific product names, identifiers, or technical terms that require exact lexical matching alongside semantic search.

### Which models on regolo support function calling for this pipeline?

`gpt-oss-120b` supports function calling for tool orchestration and extraction, while `Qwen3-Embedding-8B` provides dense vector representations for hybrid semantic retrieval.

### What accuracy improvement does graphrag deliver over vector search?

Published GraphRAG results report win rates rather than a single accuracy delta. Edge and colleagues (2024) measure 72 percent comprehensiveness and 62 percent diversity win rates for root-level GraphRAG over vector RAG, alongside 26 to 33 percent fewer context tokens for low-level community summaries and over 97 percent fewer for root-level. Treat any single headline accuracy percentage attached to graph retrieval with suspicion, including figures in vendor case studies.

---

## Ship Private AI. Not Infrastructure.

You have the private AI App architecture, bow give it an inference layer built for production.

**Regolo** gives European teams fast, OpenAI-compatible access to Mistral, Llama, Qwen, DeepSeek, GLM, and more — with zero data retention, EU data residency, and no new SDK to learn.

Change your `base_url`. Keep your LangChain code. Start shipping.

### 🚀 [Start your 30-day free trial →](https://regolo.ai/?utm_source=blog&utm_medium=cta&utm_campaign=private-rag)

Build, test, and deploy with no infrastructure to maintain.
**No credit card. No migration project. No compromise on data control.**

### 💬 [Join the Regolo Discord →](https://discord.gg/bqGrVJHeF)

Meet builders working on private RAG, local LLMs, LangChain, Ollama, and production AI systems. Share your setup, get feedback from the community, and speak directly with the Regolo team.

### 🤝 [Talk to an AI Infrastructure Engineer →](https://regolo.ai/contact?utm_source=blog&utm_medium=cta&utm_campaign=private-rag)

Running a sensitive workload, scaling beyond a proof of concept, or assessing a managed EU inference provider? Get a tailored architecture and commercial proposal for your team.

### 📂 [Clone the GitHub repository →](https://github.com/regolo-ai/tutorials/)

Get the full implementation from this guide: ingestion scripts, ChromaDB setup, hybrid retrieval, the **30-Question RAG Floor**, evaluation examples, and deployment configuration.

> **Private AI should not require a private data center.**
> Regolo gives your team an EU-native path from local experimentation to production-grade inference.

---

### Build with Regolo

- **Discord:** [Join the community →](https://discord.gg/bqGrVJHeF)
- **GitHub:** [Explore open-source workflows →](https://github.com/regolo-ai/tutorials/)
- **X / Twitter:** [Follow @regolo\_ai →](https://x.com/regolo_ai)
- **Reddit:** [Join the community →](https://www.reddit.com/r/regolo_ai/)
- **Documentation:** [Read the API docs →](https://docs.regolo.ai)
- **Contact:** [Talk to the team →](https://regolo.ai/contact)

---

*Built with ❤️ by the Regolo team. Questions? [regolo.ai/contact](https://regolo.ai/contact)* or chat with us on [Discord](https://discord.gg/bqGrVJHeF)