Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,16 @@ RUN mkdir /skaha
COPY src/startup.sh /skaha/startup.sh
RUN chmod +x /skaha/startup.sh
COPY src/launch_firefly_on_canfar.py /skaha/launch_firefly_on_canfar.py
COPY src/populate_discovery.py /skaha/populate_discovery.py
RUN chmod +x /skaha/populate_discovery.py
COPY src/jupyter /usr/bin/jupyter
# cadc_dataset_map.yaml is NOT copied into the image; it is deployed to
# /arc/projects/LSST/ via Argo CD (see arc-projects-LSST/README.md).

# Writable Nublado discovery/token paths used by lsst.rsp.RSPDiscovery.
# Skaha sessions run as arbitrary UIDs, so these must be world-writable.
RUN mkdir -p /etc/nublado/discovery /etc/nublado/secrets \
&& chmod 1777 /etc/nublado /etc/nublado/discovery /etc/nublado/secrets

# Some items to make this a better experience on CANFAR
RUN /skaha/startup.sh pip install cadctap cadcdata vos canfar safir rsp-jupyter-extensions lsdb
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,93 @@
# ARC project files for LSST on CANFAR

Files in this directory are **not** baked into the LSST science-platform
container image. They are the source of truth for configuration that must
live on shared ARC storage at:

```text
/arc/projects/LSST/
```

Session startup scripts (for example `startup.sh` in the sciplat image)
read these paths at runtime so operators can update service mappings
without rebuilding images.

## Files

| File in this directory | Deployed path | Used by |
| --- | --- | --- |
| `cadc_dataset_map.yaml` | `/arc/projects/LSST/cadc_dataset_map.yaml` | `populate_discovery.py` → `/etc/nublado/discovery/v1.json` for `lsst.rsp.RSPDiscovery` |
| `cadc_repositories.yaml` | `/arc/projects/LSST/cadc_repositories.yaml` | `DAF_BUTLER_REPOSITORY_INDEX` in `startup.sh` (also at `https://www.canfar.net/storage/arc/file/projects/LSST/cadc_repositories.yaml`) |

## Argo CD deployment

These files should be synced by the platform Argo CD configuration in
[opencadc/science-platform](https://github.com/opencadc/science-platform)
(or the LSST project’s ARC provisioning workflow), **not** by the Docker
build for this image.

Recommended pattern:

1. Keep the canonical copies in this directory (this repo).
2. Add an Argo CD / GitOps sync that publishes them to
`/arc/projects/LSST/` on the CANFAR storage backend used by Skaha
sessions.
3. After changing `cadc_dataset_map.yaml`, new notebook sessions pick up
the update on next start (discovery JSON is regenerated each startup).
4. After changing `cadc_repositories.yaml`, new sessions see updated Butler
labels via `DAF_BUTLER_REPOSITORY_INDEX`.

Do **not** `COPY` these files into the Dockerfile. The container only
ships `populate_discovery.py` and env wiring that expect the ARC paths.

## Local testing

```bash
python science-containers/Dockerfiles/lsst-science-platform/src/populate_discovery.py \
--map science-containers/Dockerfiles/lsst-science-platform/arc-projects-LSST/cadc_dataset_map.yaml \
--dry-run
```

## Dataset map notes

- Dataset labels (`dp1`, `dp2`, `dp02`, `prompt`, …) are arguments to
`RSPDiscovery("dp1")`.
- Service values may be:
- a CADC IVOID from
[resource-caps](https://ws.cadc-ccda.hia-iha.nrc-cnrc.gc.ca/reg/resource-caps)
- a literal `https://...` URL (Rubin mirrors)
- a `{url, versions}` dict
- `dp1` / `dp2`: CANFAR YouCAT, GMS, SIA, SODA cutout, DataLink, plus
Butler configs on `ws-uv.canfar.net`. HiPS points at Rubin.
- `dp02` / `prompt`: discovery entries that point at `data.lsst.cloud`.

### Dual-token authentication

`RSPDiscovery` sends **one** bearer token to every service URL in a dataset.
CANFAR (CADC) and Rubin (Gafaelfawr) tokens are not interchangeable.

| Dataset | Default session token (CADC) | Rubin token required |
| --- | --- | --- |
| `dp1`, `dp2` | tap, gms, sia, cutout, datalink | hips (Rubin URL) |
| `dp02`, `prompt` | will not work | all services |

```python
import os
from lsst.rsp import RSPDiscovery

# CANFAR services (session token from /etc/nublado/secrets/token)
canfar = RSPDiscovery("dp2")
tap = canfar.get_tap_client()

# Rubin-mirrored discovery (explicit Rubin token)
rubin = RSPDiscovery("dp02", token=os.environ["RUBIN_TOKEN"])
hips_url = rubin.get_service_url("hips")
```

Create Rubin tokens via the RSP token UI:
https://rsp.lsst.io/guides/auth/creating-user-tokens.html

## Butler repository index notes

- Keys are Butler repository labels (`dp1`, `canfar-dp1`, `rubin-dp1`, …).
- Values are Butler configuration YAML URLs for CANFAR or Rubin endpoints.
Original file line number Diff line number Diff line change
@@ -0,0 +1,140 @@
# Map RSP dataset labels to CADC/IVOA resources and optional literal URLs.
#
# Deployed location (outside the container image):
# /arc/projects/LSST/cadc_dataset_map.yaml
#
# At session start, populate_discovery.py reads this file and writes
# /etc/nublado/discovery/v1.json for lsst.rsp.RSPDiscovery.
#
# Service values may be:
# - ivo://... → resolve via resource-caps + VOSI capabilities (CANFAR/CADC)
# - https://... → literal URL (used to mirror Rubin services into discovery)
# - {url, versions} dict for fully specified literal entries
#
# AUTH WARNING
# ------------
# RSPDiscovery attaches a single bearer token to every service URL for a
# dataset. CANFAR (CADC) tokens and Rubin (Gafaelfawr) tokens are different.
# - dp1 / dp2: default session token is the CADC token from startup.sh.
# CANFAR-resolved services (tap, gms, sia, cutout, datalink) work with it.
# hips entries below point at Rubin and will NOT authenticate with that
# CADC token; call them only with an explicit Rubin token, or use the
# Rubin-mirrored datasets (dp02 / prompt) instead.
# - dp02 / prompt: URLs point at data.lsst.cloud. Use a Rubin token:
# RSPDiscovery("dp02", token=os.environ["RUBIN_TOKEN"])
# Do not expect the default CANFAR ACCESS_TOKEN / NUBLADO_TOKEN to work.
#
# Source of truth: science-containers/.../arc-projects-LSST/
# Sync to ARC via Argo CD (see README.md in that directory).

resource_caps_url: https://ws.cadc-ccda.hia-iha.nrc-cnrc.gc.ca/reg/resource-caps
discovery_path: /etc/nublado/discovery/v1.json
environment_name: canfar

# RSPDiscovery service name → IVOA capability standardID(s) to match
# inside a resource's VOSI capabilities document (IVOID entries only).
service_standard_ids:
tap:
- ivo://ivoa.net/std/TAP
sia:
- ivo://ivoa.net/std/SIA#query-2.0
- ivo://ivoa.net/std/SIA
datalink:
- ivo://ivoa.net/std/DataLink#links-1.1
- ivo://ivoa.net/std/DataLink
cutout:
- ivo://ivoa.net/std/SODA#sync-1.0
- ivo://ivoa.net/std/SODA
gms:
- ivo://ivoa.net/std/GMS#search-1.0
- ivo://ivoa.net/std/GMS#search-0.1

datasets:
# --- CANFAR-authenticated datasets (default session CADC token) ---
dp1:
description: >-
Data Preview 1 on CANFAR (YouCAT + CAOM ops + CADC GMS). HiPS URL is
Rubin's and needs a separate Rubin token.
docs_url: https://dp1.lsst.io/
butler_config: https://ws-uv.canfar.net/lsst/api/butler/repo/dp1/butler.yaml
services:
tap: ivo://cadc.nrc.ca/youcat
gms: ivo://cadc.nrc.ca/gms
sia: ivo://cadc.nrc.ca/sia
cutout: ivo://cadc.nrc.ca/caom2ops
datalink: ivo://cadc.nrc.ca/caom2ops
hips:
url: https://data.lsst.cloud/api/hips/v2/dp1/list
versions:
hips-list-1.0:
url: https://data.lsst.cloud/api/hips/v2/dp1/list

dp2:
description: >-
Data Preview 2 on CANFAR (YouCAT + CAOM ops + CADC GMS). HiPS URL is
Rubin's and needs a separate Rubin token.
docs_url: https://dp2.lsst.io/
butler_config: https://ws-uv.canfar.net/lsst/api/butler/repo/dp2/butler.yaml
services:
tap: ivo://cadc.nrc.ca/youcat
gms: ivo://cadc.nrc.ca/gms
sia: ivo://cadc.nrc.ca/sia
cutout: ivo://cadc.nrc.ca/caom2ops
datalink: ivo://cadc.nrc.ca/caom2ops
hips:
url: https://data.lsst.cloud/api/hips/v2/dp2/list
versions:
hips-list-1.0:
url: https://data.lsst.cloud/api/hips/v2/dp2/list

# --- Rubin-mirrored datasets (require a Rubin Gafaelfawr token) ---
dp02:
description: >-
DP0.2 discovery entries pointing at Rubin data.lsst.cloud services.
Authenticate with a Rubin token, not the CANFAR session token.
docs_url: https://dp0-2.lsst.io/
butler_config: https://data.lsst.cloud/api/butler/repo/dp02/butler.yaml
services:
tap:
url: https://data.lsst.cloud/api/tap
versions:
tables:
url: https://data.lsst.cloud/api/tap/tables
sia:
url: https://data.lsst.cloud/api/sia/dp02
versions:
sia-query-2.0:
url: https://data.lsst.cloud/api/sia/dp02/query
cutout:
url: https://data.lsst.cloud/api/cutout
versions:
soda-async-1.0:
url: https://data.lsst.cloud/api/cutout/jobs
soda-sync-1.0:
url: https://data.lsst.cloud/api/cutout/sync
datalink:
url: https://data.lsst.cloud/api/datalink
versions:
datalink-links-1.1:
url: https://data.lsst.cloud/api/datalink/links
gms:
url: https://data.lsst.cloud/auth/gms
versions:
gms-search-1.0:
url: https://data.lsst.cloud/auth/gms
hips:
url: https://data.lsst.cloud/api/hips/v2/dp02/list
versions:
hips-list-1.0:
url: https://data.lsst.cloud/api/hips/v2/dp02/list

prompt:
description: >-
Rubin prompt-product discovery (alerts). Requires a Rubin token.
services:
alerts: https://data.lsst.cloud/api/alerts
gms:
url: https://data.lsst.cloud/auth/gms
versions:
gms-search-1.0:
url: https://data.lsst.cloud/auth/gms
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
# Butler repository index for CANFAR LSST sessions.
#
# Deployed location (outside the container image):
# /arc/projects/LSST/cadc_repositories.yaml
#
# Referenced by startup.sh as DAF_BUTLER_REPOSITORY_INDEX
# (also served at
# https://www.canfar.net/storage/arc/file/projects/LSST/cadc_repositories.yaml).
#
# Source of truth: science-containers/.../arc-projects-LSST/
# Sync to ARC via Argo CD (see README.md in this directory).

dp1: "https://ws-uv.canfar.net/lsst/api/butler/repo/dp1/butler.yaml"
dp2: "https://ws-uv.canfar.net/lsst/api/butler/repo/dp2/butler.yaml"
canfar-dp1: "https://ws-uv.canfar.net/lsst/api/butler/repo/dp1/butler.yaml"
canfar-dp2: "https://ws-uv.canfar.net/lsst/api/butler/repo/dp2/butler.yaml"
rubin-dp1: "https://data.lsst.cloud/api/butler/repo/dp1/butler.yaml"
rubin-dp2: "https://data.lsst.cloud/api/butler/repo/dp2/butler.yaml"
Loading
Loading