Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions .github/workflows/update_dev_docs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,16 @@ jobs:
python-version: '3.10'
- name: Install dependencies
run: uv sync --locked --group dev --group docs
- name: Set up Node.js
uses: actions/setup-node@v4
with:
node-version: "20"
- name: Build documentation widgets
run: |
npm --prefix diagram-app ci
npm --prefix diagram-app run build
npm --prefix qc-tree-app ci
npm --prefix qc-tree-app run build
- name: Generate docs
run: |
uv run python src/biodata_schema/utils/docs/model_generator.py
Expand Down
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -147,10 +147,14 @@ examples/*.json
!diagram-app/package.json
!diagram-app/package-lock.json
!diagram-app/tsconfig.json
!qc-tree-app/package.json
!qc-tree-app/package-lock.json
!qc-tree-app/tsconfig.json

# diagram-app (React Flow schema diagram embedded in the docs front page)
diagram-app/node_modules/
docs/source/_static/schema-diagram/
docs/source/_static/qc-tree-app/
docs/base/models/*
.github/copilot-instructions.md
schema_tree.md
Expand Down
2 changes: 2 additions & 0 deletions .readthedocs.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,8 @@ build:
pre_build:
- npm --prefix diagram-app ci
- npm --prefix diagram-app run build
- npm --prefix qc-tree-app ci
- npm --prefix qc-tree-app run build

python:
install:
Expand Down
7 changes: 7 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,13 @@ npm --prefix diagram-app ci
npm --prefix diagram-app run build
```

The quality-control page embeds the interactive hierarchy builder (`qc-tree-app/`). Build its JS/CSS bundle before building the docs as well:

```bash
npm --prefix qc-tree-app ci
npm --prefix qc-tree-app run build
```

Then to create the documentation html files, run:

```bash
Expand Down
43 changes: 35 additions & 8 deletions docs/base/core/quality_control.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,21 +10,26 @@ Every [QCMetric](#qcmetric) has a `Status` which takes the value of the metric a

## Details

### Metrics
The metrics defined during quality control should define whether or not an asset can be used for analysis and the properties of an asset that could influence how an analysis is performed. There are two levels of QC:

Each [QCMetric](#qcmetric) is a single value or array of values that can be computed, or observed, about one modality in a data asset. These can have any type. Metrics should be significant: i.e. whether they pass or fail should matter for the modality. Metrics need to be human understandable. If you find yourself generating more than fifty metrics for a modality you should group them together (i.e. make the value a dictionary combining similar metrics and the rule an evaluation of multiple fields in the dictionary).
- Quality control metrics that are **not allowed to fail** are metrics that at `stage:raw` would prevent an asset from being processed or that at `stage:processing` would prevent an asset from being analyzed. This is a very high bar. If the goals of the data acquisition were met (regardless of whether *experimental* goals were met), then all metrics that are 'not allowed to fail' should be passing.
- All other metrics should be **allowed to fail**.

Each [QCMetric](#qcmetric) has a [Status](#status). The [Status](#status) should depend directly on the `QCMetric.value`, either by a simple function: "value>5", or by a qualitative rule: "Field of view includes visual areas". The `QCMetric.description` field should describe the rule used to set the status. Metrics can be evaluated multiple times, in which case the new status should be appended the `QCMetric.status_history`.
For example, if the goal of an acquisition was to collect behavior and fiber photometry, the 'not allowed to fail' metrics should define whether the behavior and fiber photometry data can be analyzed. Metrics that are allowed to fail could include whether the data meets various thresholds of behavior engagement, the signal to noise ratio of the data, and other potentially valuable (but not critical) metrics.

Each [QCMetric](#qcmetric) is annotated with three pieces of additional metadata: the [Stage](#stage) during which it was evaluated, the [Modality](biodata_models/modalities.md#modality) of the evaluated data, and [tags](#tags).
Metrics that *can fail* are identified by the `QualityControl.allow_tag_failures` field, details below.

### Curations
### Metrics

Each [QCMetric](#qcmetric) is a single value or array of values that can be computed, or observed, about one modality in a data asset. These can have any type. Metrics should be significant: i.e. whether they pass or fail should matter for the modality. Metrics need to be human understandable. If you find yourself generating more metrics than a human can reasonably parse for a modality you should group them together (i.e. make the value a dictionary combining similar metrics).

If you find yourself computing a value for something smaller than an entire modality of data in an asset you are performing *curation*, i.e. you are determining the status of a subset of a modality in the data asset. We provide the [CurationMetric](#curationmetric) for this purpose. You should put a dictionary in the `CurationMetric.value` field that contains a mapping between the subsets (usually neurons, ROIs, channels, etc) and their values.
Each [QCMetric](#qcmetric) has a [Status](#status). The [Status](#status) should depend directly on the `QCMetric.value`, either by a simple function: "value>5", or by a qualitative rule: "Field of view includes visual areas". The `QCMetric.description` field should describe the rule used to set the status. The status of a metric should be appended the `QCMetric.status_history`.

Each [QCMetric](#qcmetric) is annotated with three pieces of additional metadata: the [Stage](#stage) during which it was evaluated, the [Modality](biodata_models/modalities.md#modality) of the evaluated data, and [tags](#tags).

### Tags

`tags` are groups of descriptors that define how metrics are organized hierarchically, making it easier to visualize metrics. Good tag keys (groups) are things like "probe" and good tag values are things like "Probe A" or just "A".
`tags` are groups of descriptors that define how metrics are grouped hierarchically, making it easier to visualize metrics. Good tag keys (groups) are things like "probe" and good tag values (group members) are things like "Probe A" or just "A".

```python
# For an electrophysiology metric
Expand All @@ -39,7 +44,19 @@ tags = {
}
```

Use the `QualityControl.default_grouping` list to define how users should organize a visualization by default. In almost all cases *modality should be the top-level grouping*. For example, building on the example above you might group by: `["modality", ("probe", "video"), "shank"]` to get a tree split by modality first (which naturally splits ephys and behavior-videos tags into two groups), then by which probe or video a metric belongs to, and finally only for probes the individual shanks are split into groups.
When multiple QC stages are selected, they split at the top of the hierarchy. Multiple modalities split at the next level, or at the top when only one stage is selected. These fixed levels are followed by the tag levels in `QualityControl.default_grouping`. For example, `["stage", "modality", ("probe", "video"), "shank"]` groups first by stage, then by modality, then by probe or video at the same level, and finally by shank.

Use the builder to define tag keys and values, drag tag keys into level buckets, and preview the resulting hierarchy:

```{raw} html
<link rel="stylesheet" href="_static/qc-tree-app/qc-tree-app.css">
<div class="qc-tree-app"></div>
<script type="module" src="_static/qc-tree-app/qc-tree-app.js"></script>
```

### Curations

If you find yourself computing a value for something that is smaller than an entire modality of data and is repeated (e.g. neurons) in an asset you are performing *curation*, i.e. you are determining the status of a subset of a modality in the data asset. We provide the [CurationMetric](#curationmetric) for this purpose. You should put a dictionary in the `CurationMetric.value` field that contains a mapping between the subsets (usually neurons, ROIs, channels, etc) and their values.

### QualityControl.evaluate_status()

Expand All @@ -53,6 +70,16 @@ Then, given the status of all the remaining metrics in the group:
2. If any metric is pending and the rest pass the evaluation is pending
3. If all metrics pass the evaluation passes

### QualityControl.status

The `QualityControl.status` field is a dictionary that maps individual tag `key:value` pairs to their status. For example a typical status dictionary might look like this:

```
{
"modality:behavior": "PASS",
}
```

**Q: What is a metric reference?**

Each [QCMetric](#qcmetric) should include a `QCMetric.reference`. References should be publicly accessible images, figures, multi-panel figures, and videos that support the metric value/status or provide the information necessary for manual annotation.
Expand Down
43 changes: 35 additions & 8 deletions docs/source/quality_control.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,21 +10,26 @@ Every [QCMetric](#qcmetric) has a `Status` which takes the value of the metric a

## Details

### Metrics
The metrics defined during quality control should define whether or not an asset can be used for analysis and the properties of an asset that could influence how an analysis is performed. There are two levels of QC:

Each [QCMetric](#qcmetric) is a single value or array of values that can be computed, or observed, about one modality in a data asset. These can have any type. Metrics should be significant: i.e. whether they pass or fail should matter for the modality. Metrics need to be human understandable. If you find yourself generating more than fifty metrics for a modality you should group them together (i.e. make the value a dictionary combining similar metrics and the rule an evaluation of multiple fields in the dictionary).
- Quality control metrics that are **not allowed to fail** are metrics that at `stage:raw` would prevent an asset from being processed or that at `stage:processing` would prevent an asset from being analyzed. This is a very high bar. If the goals of the data acquisition were met (regardless of whether *experimental* goals were met), then all metrics that are 'not allowed to fail' should be passing.
- All other metrics should be **allowed to fail**.

Each [QCMetric](#qcmetric) has a [Status](#status). The [Status](#status) should depend directly on the `QCMetric.value`, either by a simple function: "value>5", or by a qualitative rule: "Field of view includes visual areas". The `QCMetric.description` field should describe the rule used to set the status. Metrics can be evaluated multiple times, in which case the new status should be appended the `QCMetric.status_history`.
For example, if the goal of an acquisition was to collect behavior and fiber photometry, the 'not allowed to fail' metrics should define whether the behavior and fiber photometry data can be analyzed. Metrics that are allowed to fail could include whether the data meets various thresholds of behavior engagement, the signal to noise ratio of the data, and other potentially valuable (but not critical) metrics.

Each [QCMetric](#qcmetric) is annotated with three pieces of additional metadata: the [Stage](#stage) during which it was evaluated, the [Modality](biodata_models/modalities.md#modality) of the evaluated data, and [tags](#tags).
Metrics that *can fail* are identified by the `QualityControl.allow_tag_failures` field, details below.

### Curations
### Metrics

Each [QCMetric](#qcmetric) is a single value or array of values that can be computed, or observed, about one modality in a data asset. These can have any type. Metrics should be significant: i.e. whether they pass or fail should matter for the modality. Metrics need to be human understandable. If you find yourself generating more metrics than a human can reasonably parse for a modality you should group them together (i.e. make the value a dictionary combining similar metrics).

If you find yourself computing a value for something smaller than an entire modality of data in an asset you are performing *curation*, i.e. you are determining the status of a subset of a modality in the data asset. We provide the [CurationMetric](#curationmetric) for this purpose. You should put a dictionary in the `CurationMetric.value` field that contains a mapping between the subsets (usually neurons, ROIs, channels, etc) and their values.
Each [QCMetric](#qcmetric) has a [Status](#status). The [Status](#status) should depend directly on the `QCMetric.value`, either by a simple function: "value>5", or by a qualitative rule: "Field of view includes visual areas". The `QCMetric.description` field should describe the rule used to set the status. The status of a metric should be appended the `QCMetric.status_history`.

Each [QCMetric](#qcmetric) is annotated with three pieces of additional metadata: the [Stage](#stage) during which it was evaluated, the [Modality](biodata_models/modalities.md#modality) of the evaluated data, and [tags](#tags).

### Tags

`tags` are groups of descriptors that define how metrics are organized hierarchically, making it easier to visualize metrics. Good tag keys (groups) are things like "probe" and good tag values are things like "Probe A" or just "A".
`tags` are groups of descriptors that define how metrics are grouped hierarchically, making it easier to visualize metrics. Good tag keys (groups) are things like "probe" and good tag values (group members) are things like "Probe A" or just "A".

```python
# For an electrophysiology metric
Expand All @@ -39,7 +44,19 @@ tags = {
}
```

Use the `QualityControl.default_grouping` list to define how users should organize a visualization by default. In almost all cases *modality should be the top-level grouping*. For example, building on the example above you might group by: `["modality", ("probe", "video"), "shank"]` to get a tree split by modality first (which naturally splits ephys and behavior-videos tags into two groups), then by which probe or video a metric belongs to, and finally only for probes the individual shanks are split into groups.
When multiple QC stages are selected, they split at the top of the hierarchy. Multiple modalities split at the next level, or at the top when only one stage is selected. These fixed levels are followed by the tag levels in `QualityControl.default_grouping`. For example, `["stage", "modality", ("probe", "video"), "shank"]` groups first by stage, then by modality, then by probe or video at the same level, and finally by shank.

Use the builder to define tag keys and values, drag tag keys into level buckets, and preview the resulting hierarchy:

```{raw} html
<link rel="stylesheet" href="_static/qc-tree-app/qc-tree-app.css">
<div class="qc-tree-app"></div>
<script type="module" src="_static/qc-tree-app/qc-tree-app.js"></script>
```

### Curations

If you find yourself computing a value for something that is smaller than an entire modality of data and is repeated (e.g. neurons) in an asset you are performing *curation*, i.e. you are determining the status of a subset of a modality in the data asset. We provide the [CurationMetric](#curationmetric) for this purpose. You should put a dictionary in the `CurationMetric.value` field that contains a mapping between the subsets (usually neurons, ROIs, channels, etc) and their values.

### QualityControl.evaluate_status()

Expand All @@ -53,6 +70,16 @@ Then, given the status of all the remaining metrics in the group:
2. If any metric is pending and the rest pass the evaluation is pending
3. If all metrics pass the evaluation passes

### QualityControl.status

The `QualityControl.status` field is a dictionary that maps individual tag `key:value` pairs to their status. For example a typical status dictionary might look like this:

```
{
"modality:behavior": "PASS",
}
```

**Q: What is a metric reference?**

Each [QCMetric](#qcmetric) should include a `QCMetric.reference`. References should be publicly accessible images, figures, multi-panel figures, and videos that support the metric value/status or provide the information necessary for manual annotation.
Expand Down
2 changes: 2 additions & 0 deletions qc-tree-app/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
node_modules/
dist/
17 changes: 17 additions & 0 deletions qc-tree-app/index.html
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="UTF-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<title>QC Tree Builder</title>
<style>
html, body { min-height: 100%; margin: 0; }
body { color: #111; background: #fff; font: 16px Arial, Helvetica, sans-serif; }
#root { box-sizing: border-box; width: min(100%, 560px); margin: 0 auto; padding: 24px 12px; }
</style>
</head>
<body>
<div id="root"></div>
<script type="module" src="/src/main.tsx"></script>
</body>
</html>
Loading
Loading