The HCA Data Browser is built using Next.js.
Node.js 22.12.0 is required to run the app.
git clone https://github.com/DataBiosphere/data-browser.git [folder_name]
From the project root directory, install client-side dependencies:
npm install
To start the development server, run the following from the explorer directory:
npm run dev:hca-dcp
You can hit the server at http://localhost:3000.
This project has end-to-end tests powered by Playwright, currently only for the anvil-cmg configuration and in progress for anvil-catalog. To run tests, run npm run test:anvil-cmg from the explorer folder. Tests will also run by default on pull request.
When updating tabs and columns on the anvil-cmg configuration, please update explorer/e2e/anvil/anvil-tabs.ts to reflect the changes
To update HCA scripts in the HCA Data Explorer, navigate to the explorer directory and run:
npm run get-cellxgene-projects-hcaThis will save any updates to explorer/site-config/hca-dcp/ma-dev/scripts/out/cellxgene-projects.json based on HCA links provided by CELLxGENE.
The Diagnosis and Phenotype filters and columns in the AnVIL Data Explorer show names for HP, OMIM and Orphanet term IDs, for example "Abnormality of the kidney (HP:0000077)". The names come from site-config/anvil-cmg/dev/index/common/diagnosis.ts, which is generated. An ID that isn't in that file shows up raw.
- New AnVIL datasets have been indexed, or raw IDs such as
OMIM:310200show up in the Diagnosis filter or column. - The Human Phenotype Ontology or Orphadata has published a new release.
Run:
npm run refresh-diagnosis-terms:anvil-cmgNo login is needed. The script:
- Collects the term IDs in the
diagnoses.diseaseanddiagnoses.phenotypefacets of the public AnVIL Azul datasets endpoint, and keeps the IDs already indiagnosis.ts. - Looks up every ID in the latest source files, so names that have changed upstream are updated:
- HP:
hp.obofrom the latest Human Phenotype Ontology release - OMIM:
phenotype.hpoafrom the same release - Orphanet (
ORPHA:orOrphanet:): Orphadata's disease list,en_product1.xml
- HP:
- Writes
diagnosis.ts, formatted with Prettier.
It takes about a minute, most of it downloading the source files.
The last lines list the IDs that have no name, grouped by prefix. They will show up raw in the UI.
- An HP, OMIM or Orphanet ID in this list isn't in the source files. That's expected for a handful of retired OMIM IDs.
- Other prefixes, such as
MONDO, are not looked up. Malformed values, such asH:0010609, are data problems to report to the Azul team.
The script stops with an error, and leaves diagnosis.ts unchanged, if:
- AnVIL Azul is indexing or down. Its facets may be incomplete while it indexes; run the script again once indexing has finished.
- the Azul response has neither the
diagnoses.diseasenor thediagnoses.phenotypefacet. The Azul API has probably changed; check the response before retrying. - a download takes longer than 2 minutes. This is usually a stalled connection; run the script again.
- a source file gives no names at all. The download probably returned an error page, or the file's format has changed. Open the URL in the error to see which.
- Run
git diff site-config/anvil-cmg/dev/index/common/diagnosis.ts. Expect added IDs and a few updated names. An ID is removed only if the sources no longer name it, and it then appears in the list of IDs with no name. - Run
npm run dev:anvil-cmg, open the Diagnosis filter, and check that the IDs you were fixing now show names. - Commit
diagnosis.ts.
The script reads the datasets endpoint only, because the other endpoints return diagnosis values only to signed-in users, and Azul keeps at most 100 diagnosis values per dataset (azul#8369). An ID that appears only on the Donors or BioSamples tabs, and isn't already in diagnosis.ts, stays raw until it shows up in the datasets facets.