REST API for serving INSEE files v3 with full-text search and geographic search capabilities.
To have a working copy of this project, follow the instructions.
-
Setup Rust: Install Rust (version 1.70+ recommended)
-
Environment variables: Define your environment variables as defined in
.env.sample. You can either manually define these environment variables or use a.envfile. -
PostgreSQL database: Setup PostgreSQL with required extensions (macOS commands):
brew install postgresql
createuser --pwprompt sirene # set password to sirenepw for instance
createdb --owner=sirene sirene
# Connect to database and enable required extensions
psql -U sirene -d sirene
CREATE EXTENSION IF NOT EXISTS postgis;
CREATE EXTENSION IF NOT EXISTS pg_trgm;
CREATE EXTENSION IF NOT EXISTS unaccent;
\q-
Required PostgreSQL extensions:
postgis(for geographic search)pg_trgm(for full-text search with trigram similarity)unaccent(for accent-insensitive search, via immutable wrapper)
-
Optional: For development, you may want to install:
brew install diesel_cli # For database migrations
cargo install cargo-watch # For auto-reloading during developmentRecommended configuration for production with docker:
RUST_LOG=sirene=warn
SIRENE_ENV=production
BASE_URL=[Your base URL, needed to update asynchronously]
API_KEY=[Any randomized string, needed to use the HTTP admin endpoint]
DATABASE_URL=postgresql://[USER]:[PASSWORD]@[PG_HOST]:[PG_PORT]/[PG_DATABASE]
DATABASE_POOL_SIZE=100
INSEE_CREDENTIALS=[API_KEY]
How to generate INSEE_CREDENTIALS
This variable is only needed if you want to have the daily updates.
- Go to https://portail-api.insee.fr/catalog/all
- Create an account or sign in
- Create an application on this portal
- Subscribe this application to API SIRENE (Sirene 4 - v3.11)
- Generate a key in the application details
- Copy the key paste it in
.envinstead of[API_KEY]
> sirene --help
Sirene service used to update data in database and serve it through a HTTP REST API
Usage: sirene <COMMAND>
Commands:
update Update data from CSV source files
serve Serve data from database to /unites_legales/<siren> and /etablissements/<siret>
help Print this message or the help of the given subcommand(s)
Options:
-h, --help Print help
-V, --version Print version
> sirene serve --help
Serve data from database to /unites_legales/<siren> and /etablissements/<siret>
Usage: sirene serve [OPTIONS] --env <ENVIRONMENT> --port <PORT> --host <HOST>
Options:
--env <ENVIRONMENT> Configure log level [env: SIRENE_ENV=development] [possible values: development, staging, production]
--port <PORT> Listen this port [env: PORT=3000]
--host <HOST> Listen this host [env: HOST=localhost]
--api-key <API_KEY> API key needed to allow maintenance operation from HTTP [env: API_KEY=]
--base-url <BASE_URL> Base URL needed to configure asynchronous polling for updates [env: BASE_URL=http://localhost:3000]
-h, --help Print help
> sirene update --help
Update data from CSV source files
Usage: sirene update [OPTIONS] <GROUP_TYPE> [COMMAND]
Commands:
update-data Download, unzip and load CSV file in database in loader-table
swap-data Swap loader-table to production
sync-insee Synchronise daily data from INSEE since the last modification
finish-error Set a staled update process to error, use only if the process is really stopped
help Print this message or the help of the given subcommand(s)
Arguments:
<GROUP_TYPE> Configure which part will be updated [possible values: unites-legales, etablissements, all]
Options:
--force Force update even if the source data where not updated
-h, --help Print help
GET /v3/unites_legales/<siren>
GET /v3/etablissements/<siret>
Search Establishments
GET /v3/etablissements?q=<text>&lat=<latitude>&lng=<longitude>&radius=<meters>&sort=<field>&direction=<asc|desc>&limit=<number>&offset=<number>
Search Legal Units
GET /v3/unites_legales?q=<text>&sort=<field>&direction=<asc|desc>&limit=<number>&offset=<number>
Query Parameters:
q: Full-text search query (searches in denomination and commune name for establishments, denomination only for legal units)lat,lng,radius: Geographic search (establishments only) - filters results within radius meters from (lat,lng) pointsort: Sort field -distance(geo only),relevance(text search),date_creation,date_debutdirection: Sort direction -ascordesc(defaults to sensible values per sort field)limit: Results per page (default: 20, max: 100)offset: Pagination offset (default: 0, max: 10000)
totalfield: exact count for filter-only queries; capped at 10,000 for text searches (q) — iftotal == 10000there may be more results.
etat_administratif: Filter by administrative status (A=active, F=closed)code_postal: Filter by postal codesiren: Filter by SIREN (establishments only)code_commune: Filter by commune codeactivite_principale: Filter by main activity codeetablissement_siege: Filter by headquarters status (establishments only)categorie_juridique: Filter by legal category (legal units only)categorie_entreprise: Filter by company category (legal units only)date_creation: Filter by creation date (legal units only)date_debut: Filter by start date (legal units only)
Maintenance
This API is enabled only if you have provided an API_KEY when starting the serve process.
POST /admin/update
{
api_key: string,
group_type: "UnitesLegales" | "Etablissements" | "All",
force: bool,
asynchronous: bool,
}
If asynchronous is set to true, the update endpoint will immediately return the following:
Status: 202 Accepted
Location: /admin/update/status?api_key=string
Retry-After: 10
[Initial status for the started update]
GET /admin/update/status?api_key=string
If an update is in progress, the status code will be 202, otherwise 200.
POST /admin/update/status/error
{
api_key: string,
}
Serve:
cargo run serve
Update:
cargo run update all
Help:
cargo run help
- REST API for INSEE SIREN/SIRET data
- Automatic updates from INSEE API
- PostgreSQL backend with efficient indexing
- Docker support for easy deployment
- Full-text search: Trigram similarity (
pg_trgm) with case and accent-insensitive matching for partial and fuzzy matches - Geographic search: Radius filtering and distance-based sorting using PostGIS
- Field filtering: Filter by administrative status, activity codes, dates, etc.
- Flexible sorting: By relevance, distance, or dates
- Pagination: Efficient offset/limit pagination — exact total for filter-only queries, capped at 10,000 for text searches (
q) to avoid full-table counts
- PostgreSQL extensions: PostGIS for spatial data, pg_trgm + unaccent for full-text search
- Optimized queries: Raw SQL with parameterized queries for performance
- OpenAPI documentation: Complete API documentation via Scalar
- Async support: Optional asynchronous updates for large datasets
cargo testA docker image is built and a sample docker-compose.yml with its docker folder are usable to test it.
docker-compose up -dRequired for production:
RUST_LOG=sirene=warn
SIRENE_ENV=production
BASE_URL=https://your-domain.com
API_KEY=your-secret-key
DATABASE_URL=postgresql://user:password@db:5432/sirene
DATABASE_POOL_SIZE=100
INSEE_CREDENTIALS=your-insee-api-key
Si vous avez une installation existante qui utilise l'extension pg_search (ParadeDB), exécutez le script de migration manuelle fourni à la racine du projet :
psql -U sirene -d sirene -f migrate_from_pg_search.sqlCe script (transactionnel) :
- Supprime les anciens index BM25 ParadeDB
- Active
pg_trgmetunaccent - Crée la fonction
immutable_unaccent()(wrapper IMMUTABLE requis pour les index) - Reconstruit
search_denominationavec normalisationlower(immutable_unaccent(...)) - Crée les nouveaux index GIN
- Recrée les tables de staging pour hériter des nouveaux index et colonnes
- Désactive l'extension
pg_search
Les nouvelles installations n'ont pas besoin de ce script : les migrations Diesel
utilisent directement pg_trgm. La migration 2026-09-15-120000_search_fts_commune
prend ensuite le relais et supprime search_denomination.
La recherche texte (?q=) repose sur le full-text search natif de PostgreSQL
(tsvector / tsquery, configuration french), sans aucune extension. Trois
mécanismes l'entourent :
- Correction de la requête.
search_lexiconcontient le vocabulaire du corpus avec sa fréquence documentaire.public.search_query(q, source)corrige les mots suspects en s'appuyant sur les trigrammes puis sur la distance de Levenshtein, et renvoie unetsqueryaugmentée : le mot saisi est conservé et complété par| correction, jamais remplacé. Un nom rare et légitime reste donc toujours trouvable. - Repli trigramme. Si le FTS ne ramène rien, la requête est rejouée une fois
avec
word_similarity. Cela couvre les transpositions et la correspondance infixe, au prix d'un index GIN supplémentaire. Ce chemin ne s'exécute jamais sur une recherche qui aboutit. - Correction phonétique. Quand le trigramme ne propose aucun candidat, le
lexique est interrogé sur une clé phonétique (
metaphoneprécédé du retrait duhinitial, muet en français). C'est ce qui reliefilipeàphilippe, que la distance d'édition seule ne rapproche pas. - Longueur minimale.
qdoit faire au moins 3 caractères, sinon l'API répond 400. En dessous, aucune structure d'index n'est exploitable.
libelle_commune ne fait pas partie du texte indexé. Les communes sont
adressées par le paramètre ?commune= en clair (paris, saint etienne,
marseile), résolu via la table de dimension commune_dim : correspondance par
préfixe sur chaque mot, repli trigramme tolérant aux fautes, puis filtrage sur
code_commune.
Les filtres de liste acceptent plusieurs valeurs séparées par des virgules, et
disposent tous d'un jumeau _not pour l'exclusion :
?code_postal=75001,75002
?activite_principale=10.71C,47.24Z
?activite_principale_not=10.71C
Une exclusion conserve les lignes dont le champ est nul : « pas 10.71C » ne dit rien des activités inconnues.
Les dates se bornent avec _min / _max, chacun facultatif et inclusif :
?date_creation_min=2024-01-01&date_creation_max=2024-12-31
?date_debut_min=2024-01-01
?facette=activite_principale,code_commune renvoie les effectifs par valeur,
calculés sur le même sous-ensemble borné que total — donc sans coût notable
(6,5 ms mesurés). Les champs autorisés sont listés dans la documentation
OpenAPI ; tout autre champ donne un 400 plutôt qu'un silence.
{
"facettes": {
"activite_principale": [
{ "valeur": "10.71C", "nombre": 7314 },
{ "valeur": "10.71A", "nombre": 380 }
]
}
}total est plafonné à 10 000 : au-delà, le compte exact coûterait un parcours
complet. total_capped vaut true quand ce plafond est atteint, pour que le
client distingue « exactement 10 000 » de « au moins 10 000 ».
Par défaut, limit (20, max 100) et offset (max 10 000). Au-delà de ce
plafond, la pagination par curseur :
GET /v3/etablissements?sort=siret&limit=1000
→ { "etablissements": [...], "next_cursor": "MDA1NTgwMTIxMDAxMDA" }
GET /v3/etablissements?sort=siret&limit=1000&cursor=MDA1NTgwMTIxMDAxMDA
next_cursor est absent sur la dernière page. Le curseur n'encode que la
dernière clé primaire rendue : il n'est valide que rejoué avec les mêmes
filtres. Un curseur illisible donne un 400, jamais un retour silencieux à la
première page.
Elle exige sort=siret (ou sort=siren) et s'exclut avec offset. Ce n'est
pas une restriction arbitraire :
- sur pertinence ou distance, la reprise par clé n'accélérerait rien — le score et la distance se recalculent ligne à ligne, il n'y a aucun parcours à raccourcir ;
- sur une date, elle exigerait un index composé : mesuré, le groupe d'ex æquo le plus dense compte 583 632 lignes et une reprise y coûte 38,8 s sans lui ;
- sur la clé primaire, l'index existe déjà et le coût par page est constant — 57 ms par page de 1 000 lignes, mesuré sur un parcours de 20 000 lignes.
C'est aussi la sémantique de l'API Insee, dont le tri par défaut est sur le siren.
Elles sont rafraîchies automatiquement par le workflow de mise à jour :
- après le swap du stock mensuel :
public.search_refresh_full(source)reconstruit le lexique etcommune_dim, puis relanceANALYZE. Sans ce dernier, la table qui vient d'être renommée n'a pas de statistiques représentatives et le planificateur repart en seq scan ; - après la synchro quotidienne Insee :
public.search_refresh_incremental(source, since)ne reparcourt que les lignes touchées depuissinceet fusionne leur vocabulaire.
Aucune des deux opérations ne bloque les lectures : mesurées sous charge, elles laissent la latence de recherche inchangée (médiane 8,9 ms pendant une reconstruction complète de 78 s, contre 8,7 ms au repos, zéro attente de verrou).
cargo test # tests unitaires seuls
SIRENE_TEST_DATABASE_URL=… cargo test # + tests d'intégrationLes tests d'intégration se mettent en sommeil sans SIRENE_TEST_DATABASE_URL ;
la variable est volontairement distincte de DATABASE_URL pour qu'un
cargo test ne puisse pas toucher la base de développement par accident.
Deux garde-fous encadrent la recherche :
- un test unitaire vérifie que l'expression construite par
src/models/search.rsfigure mot pour mot dans le DDL des index ; - un test d'intégration lit le plan d'exécution et vérifie que PostgreSQL choisit bien l'index.
Ensemble, ils ferment la porte à une recherche qui repasserait silencieusement en seq scan — le mode de panne le plus coûteux et le moins visible de cette architecture.
# Start the server
cargo run -- serve --env development --port 8080 --host 0.0.0.0
# Run tests
cargo test
# Run with auto-reload
cargo watch -x 'run -- serve --env development --port 8080'By default, serve and update apply the pending migrations on startup. To
migrate as a separate step instead (deployment hook, read-only database for the
API), run migrate beforehand and start the other commands with
--skip-migrations (or SKIP_MIGRATIONS=true): they then only check, with
read-only queries, that no migration is pending, and refuse to start otherwise.
# Apply the pending migrations, then exit
sirene migrate
# List the pending migrations, exit with 1 if there is any
sirene migrate --check
# Start without migrating (DATABASE_URL may target a read replica)
SKIP_MIGRATIONS=true sirene serve
# Run migrations with the Diesel CLI
diesel migration run
# Create new migration
diesel migration generate migration_nameThe API includes comprehensive OpenAPI documentation accessible at:
/scalar- Interactive Scalar API documentation/openapi.json- OpenAPI specification
# Text search
curl "http://localhost:8080/v3/etablissements?q=boulangerie&limit=5"
# Geographic search (within 1km of Eiffel Tower)
curl "http://localhost:8080/v3/etablissements?lat=48.8584&lng=2.2945&radius=1000&sort=distance"
# Combined search with filters
curl "http://localhost:8080/v3/etablissements?q=restaurant&code_postal=75001&etat_administratif=A&sort=relevance&limit=10"# Text search with sorting
curl "http://localhost:8080/v3/unites_legales?q=creati&sort=date_creation&direction=desc&limit=5"
# Filter by activity code
curl "http://localhost:8080/v3/unites_legales?activite_principale=62.01Z&categorie_juridique=5710"- Julien Blatecky - @Julien1619