Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,7 @@ debug_*.py

# Virtual environments
.venv/
graphrag_sdk/.venv
.venv-*/
venv/

Expand Down
68 changes: 68 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,64 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Fixed

- Re-ingesting the same file no longer duplicates its chunks. Chunk ids are
now derived from document id + position + text instead of a fresh
`uuid4()`, so `MERGE` finds the existing node (17 → 17 → 17 chunks across
three ingests, previously 17 → 34 → 51). `ContextualChunking` hashes the
original chunk text, not the LLM-enriched one. Applies to
`IngestionPipeline.run()` directly as well as through `GraphRAG.ingest()`:
when no `document_info` is supplied the pipeline now derives a stable
Document id from the source (normalised filesystem path; URIs kept
verbatim), and in text mode from a hash
of the text. `GraphRAG.ingest(text=...)` without `document_id` uses the
same text hash (previously a fresh `text-<uuid>` per call), so ingesting
identical text twice is now one document and a no-op the second time; pass
`document_id` to store identical text as distinct documents.
**Breaking:** code that called `ingest(text=...)` once per record and
relied on every call creating its own Document now gets one Document per
distinct text and a zero-result `skipped_unchanged` no-op for each
duplicate — pass a `document_id` per record to keep the old behaviour.
A caller-supplied `document_info` is merged field by field using
`model_fields_set`, so `DocumentInfo(path=...)` without an explicit `uid`
takes the derived id instead of the model's random default. Ids the
pipeline derives from a source path are checked for the reserved
`__pending__` marker (`ValueError`), the same rule `GraphRAG` applies to
explicit ids.
**Upgrade note:** graphs built before this change hold random chunk
ids; the first re-ingest of an existing document adds one more copy of its
chunk layer (matching nothing), and is stable from the second re-ingest on.
Because position is part of the id, inserting a paragraph into an edited
document re-ids every later chunk; only byte-identical files are a no-op.
Chunks written by `update()` are keyed on its transient pending id (the
pending and live documents must not share chunk nodes during the
cutover); the `content_hash` the cutover records is what keeps a later
`ingest()` of that content a no-op.
- Re-ingesting an unchanged document is now a true no-op. The pipeline hashes
the loaded text and, if the stored Document carries the same
`content_hash`, returns before chunking with
`IngestionResult.metadata["skipped_unchanged"] = True` and
`chunks_indexed = 0` — no chunker, NER, LLM or graph calls (measured: 33
provider calls and ~30 s → 0 and 0.01 s). The hash is written only after a
run completes with every write reported in full: a failed extraction, a
`RELATES`/`MENTIONED_IN` edge the graph store dropped after a transient
error, or a chunk left without an embedding by a rate-limited embedder all
withhold the hash (`IngestionResult.metadata["incomplete_writes"]` lists
the shortfalls), so the next ingest repairs the document instead of
skipping it. `update()` applies the same gate: its cutover promotes the
pending Document without a hash when the pipeline reported a shortfall
(and crash recovery of such a pending rolls forward uncertified rather
than refusing). Deployments without an embedder are not penalised —
`VectorStore.index_chunks` now returns `None` (nothing attempted) instead
of `0` (every embedding failed) when no embedder is configured, so
graph-only ingests still record the hash. Graph-store adapters without
`get_document_record` keep working — the pipeline skips the check instead
of raising.
- `GraphRAG.update(..., force=True)` re-extracts a document whose content
hash is unchanged. With `ingest()` now skipping unchanged documents, this
is the supported way to re-chunk or re-extract existing text after
changing the ontology, chunker, extractor or model (for example to adopt
the new 384-token default below); `update_sync()` takes the same flag.

- **`redis` 8.1 broke the first query on every fresh install** —
`falkordb`'s cluster probe forwarded async-pool kwargs to the sync
`redis.Redis()` constructor, which rejects them. Fixed upstream in
Expand All @@ -38,6 +96,16 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Changed

- Default chunk size lowered from 512 to 384 tokens in
`SentenceTokenCapChunking`, `StructuralChunking`, `ContextualChunking` and
the documented `CallableChunking` example. Measured on the benchmark corpus:
entity F1 0.574 vs 0.563 and relation F1 0.237 vs 0.223 against 768; with
the current extraction prompt, exact relation F1 2.2× and answer accuracy
27 → 32 % for the full stack. **Cost:** more extraction LLM calls per
ingest — 103 → 157 (+52 %) on the 11-document benchmark corpus, +35 % on a
53k-token corpus — and ~19 % more input tokens, since the per-call
instructions are re-sent once per chunk. Pass
`max_tokens=512` to keep the old size.
- Documentation migrated from MkDocs to [Mintlify](https://mintlify.com) and
published at <https://docs.falkordb.com/graphrag>, where GraphRAG SDK now
appears as a product in the FalkorDB docs product switcher. Pages moved from
Expand Down
2 changes: 1 addition & 1 deletion docs/benchmark.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -87,7 +87,7 @@ model and the judge.
| Generation temperature | 0.7 | Appendix H.2 |
| Framework | GraphRAG-SDK 1.3.0 (PyPI) on FalkorDB | — |
| Graph layout | one graph per corpus document | — |
| Chunking | `SentenceTokenCapChunking`, max_tokens 512, overlap 2 sentences | SDK default |
| Chunking | `SentenceTokenCapChunking`, max_tokens 384 (was 512 before this release), overlap 2 sentences | SDK default |
| Retrieval | `MultiPathRetrieval` — chunk_top_k 15, rel_top_k 15, max_entities 30, max_relationships 20, keyword_limit 10 | SDK default |
| Embeddings | `text-embedding-3-large` @ 1024 dimensions | Declared below |
| Text-to-Cypher | enabled | Declared below |
Expand Down
2 changes: 1 addition & 1 deletion docs/graphrag-accuracy-benchmark.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -103,7 +103,7 @@ model and the judge.
| Generation temperature | 0.7 | Appendix H.2 |
| Framework | GraphRAG-SDK 1.3.0 (PyPI) on FalkorDB | — |
| Graph layout | one graph per corpus document | — |
| Chunking | `SentenceTokenCapChunking`, max_tokens 512, overlap 2 sentences | SDK default |
| Chunking | `SentenceTokenCapChunking`, max_tokens 512, overlap 2 sentences | SDK default at 1.3.0 (the default is now 384) |
| Retrieval | `MultiPathRetrieval` — chunk_top_k 15, rel_top_k 15, max_entities 30, max_relationships 20, keyword_limit 10 | SDK default |
| Embeddings | `text-embedding-3-large` @ 1024 dimensions | Declared deviation |
| Text-to-Cypher | enabled (`enable_cypher=True`) | Declared deviation |
Expand Down
7 changes: 4 additions & 3 deletions docs/incremental-updates.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -53,11 +53,12 @@ result = await rag.update(

| Argument | Meaning |
| --------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `source` | File path. Mutually exclusive with `text`. In file mode, `document_id` defaults to `os.path.normpath(source)`. |
| `source` | File path. Mutually exclusive with `text`. In file mode, `document_id` defaults to the normalised path (a URI source is used verbatim) — the same id `ingest(source)` derives. |
| `text` | Raw text. Skips the loader. Requires an explicit `document_id`. |
| `document_id` | Stable id of the `Document` node to update. Required for text mode. |
| `loader` / `chunker` / `extractor` / `resolver` | Per-call strategy overrides, identical to `ingest()`. |
| `if_missing` | `"error"` (default) raises `DocumentNotFoundError` if the id is unknown. `"ingest"` falls through to a fresh `ingest()` — upsert semantics. |
| `force` | `False` (default) returns `no_op=True` when the content hash is unchanged. `True` re-extracts anyway — the way to apply a new ontology, chunker or extractor to existing text, since `ingest()` also skips unchanged documents. |

### Returns — `UpdateResult`

Expand All @@ -79,7 +80,7 @@ UpdateResult(

| Scenario | What changes |
| ----------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Content hash matches (touch-only edit) | **Nothing.** SHA-256 short-circuits the call to a single Cypher lookup. `no_op=True`. Use this for CRLF/formatter-only PRs. |
| Content hash matches (touch-only edit) | **Nothing.** SHA-256 short-circuits the call to a single Cypher lookup. `no_op=True`. Use this for CRLF/formatter-only PRs. Pass `force=True` to re-extract regardless. |
| Real content change | New chunks written under a pending Document, then atomically cut over to the canonical id. Old chunks are deleted. Entities previously mentioned only by this document are removed. |
| Entity still referenced by another doc | **Preserved.** Orphan cleanup is scoped to candidates from this document — never global. |
| `RELATES` edge sourced only from old chunks | **Removed** as a stale fact (cleanup keyed on `source_chunk_ids`). |
Expand Down Expand Up @@ -124,7 +125,7 @@ result = await rag.delete_document(

| Argument | Meaning |
| ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `document_id` | The Document node id (e.g. `os.path.normpath(path)` for file-mode ingests). |
| `document_id` | The Document node id (the normalised path for file-mode ingests; `text-<16hex>` for text-mode ingests without an explicit id). |
| `if_missing` | `"error"` (default) raises `DocumentNotFoundError`. `"ignore"` returns an empty result with zero counts — useful for CI deletes when the caller doesn't track which files were ever ingested. |

### Returns — `DeleteDocumentResult`
Expand Down
6 changes: 3 additions & 3 deletions docs/strategies.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -133,7 +133,7 @@ Splits at sentence boundaries (never mid-sentence) and enforces a hard token cap
from graphrag_sdk.ingestion.chunking_strategies.sentence_token_cap import SentenceTokenCapChunking

chunker = SentenceTokenCapChunking(
max_tokens=512, # max tokens per chunk (default: 512)
max_tokens=384, # max tokens per chunk (default: 384)
overlap_sentences=2, # sentences shared between chunks (default: 2)
encoding_name="cl100k_base", # tiktoken encoding (default: cl100k_base)
)
Expand All @@ -148,7 +148,7 @@ from graphrag_sdk.ingestion.chunking_strategies.contextual_chunking import Conte

chunker = ContextualChunking(
llm=my_llm,
max_tokens=512, # token cap per chunk (default: 512)
max_tokens=384, # token cap per chunk (default: 384)
overlap_sentences=2, # sentence overlap (default: 2)
max_document_tokens=16_000, # truncation limit for the doc reference in prompts (default: 16000)
)
Expand Down Expand Up @@ -177,7 +177,7 @@ Groups content by heading hierarchy into token-bounded chunks. Each chunk stores
from graphrag_sdk.ingestion.chunking_strategies.structural_chunking import StructuralChunking

chunker = StructuralChunking(
max_tokens=512, # max tokens per chunk (default: 512)
max_tokens=384, # max tokens per chunk (default: 384)
overlap_sentences=2, # sentences shared between chunks (default: 2)
)
```
Expand Down
2 changes: 1 addition & 1 deletion graphrag_sdk/examples/06_markdown_document_aware.py
Original file line number Diff line number Diff line change
Expand Up @@ -194,7 +194,7 @@ async def main():
result = await rag.ingest(
md_path,
loader=MarkdownLoader(),
chunker=StructuralChunking(max_tokens=512),
chunker=StructuralChunking(max_tokens=384),
)
print(f"Done: {result.nodes_created} nodes, {result.relationships_created} edges, "
f"{result.chunks_indexed} chunks indexed")
Expand Down
Loading
Loading