Skip to content

Commit 9754b07

Browse files
author
Andrey Cheptsov
committed
Add CLI and HTTP API guide
1 parent 6657451 commit 9754b07

58 files changed

Lines changed: 289 additions & 292 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎mkdocs.yml‎

Lines changed: 19 additions & 19 deletions
Original file line numberDiff line numberDiff line change
@@ -361,25 +361,25 @@ nav:
361361
- dstack export: docs/reference/cli/dstack/export.md
362362
- dstack import: docs/reference/cli/dstack/import.md
363363
- HTTP API:
364-
- Users: docs/reference/http/users.md
365-
- Projects: docs/reference/http/projects.md
366-
- Backends: docs/reference/http/backends.md
367-
- Fleets: docs/reference/http/fleets.md
368-
- Runs: docs/reference/http/runs.md
369-
- Gateways: docs/reference/http/gateways.md
370-
- Volumes: docs/reference/http/volumes.md
371-
- Logs: docs/reference/http/logs.md
372-
- Events: docs/reference/http/events.md
373-
- Repos: docs/reference/http/repos.md
374-
- Files: docs/reference/http/files.md
375-
- Secrets: docs/reference/http/secrets.md
376-
- Exports: docs/reference/http/exports.md
377-
- Metrics: docs/reference/http/metrics.md
378-
- GPUs: docs/reference/http/gpus.md
379-
- Templates: docs/reference/http/templates.md
380-
- Proxy: docs/reference/http/proxy.md
381-
- Authentication: docs/reference/http/authentication.md
382-
- Server: docs/reference/http/server.md
364+
- users: docs/reference/http/users.md
365+
- projects: docs/reference/http/projects.md
366+
- backends: docs/reference/http/backends.md
367+
- fleets: docs/reference/http/fleets.md
368+
- runs: docs/reference/http/runs.md
369+
- gateways: docs/reference/http/gateways.md
370+
- volumes: docs/reference/http/volumes.md
371+
- logs: docs/reference/http/logs.md
372+
- events: docs/reference/http/events.md
373+
- repos: docs/reference/http/repos.md
374+
- files: docs/reference/http/files.md
375+
- secrets: docs/reference/http/secrets.md
376+
- exports: docs/reference/http/exports.md
377+
- metrics: docs/reference/http/metrics.md
378+
- gpus: docs/reference/http/gpus.md
379+
- templates: docs/reference/http/templates.md
380+
- proxy: docs/reference/http/proxy.md
381+
- authentication: docs/reference/http/authentication.md
382+
- server: docs/reference/http/server.md
383383
- Environment variables: docs/reference/env.md
384384
- More:
385385
- .dstack/profiles.yml: docs/reference/profiles.yml.md

‎mkdocs/docs/concepts/services.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -125,7 +125,7 @@ If you do not have a [gateway](gateways.md) created, the service endpoint will b
125125
```shell
126126
$ curl http://localhost:3000/proxy/services/main/qwen36/v1/chat/completions \
127127
-H 'Content-Type: application/json' \
128-
-H 'Authorization: Bearer <dstack token>' \
128+
-H 'Authorization: Bearer <user token>' \
129129
-d '{
130130
"model": "Qwen/Qwen3.6-27B",
131131
"messages": [
@@ -143,7 +143,7 @@ The request and response format depends on the serving framework used by the
143143
service. Even for OpenAI-compatible endpoints, the format may vary slightly
144144
across frameworks.
145145

146-
If [authorization](#authorization) is not disabled, the service endpoint requires the `Authorization` header with `Bearer <dstack token>`.
146+
If [authorization](#authorization) is not disabled, the service endpoint requires the `Authorization` header with `Bearer <user token>`.
147147

148148
## Configuration options
149149

@@ -173,7 +173,7 @@ If you have a [gateway](gateways.md) created, the service endpoint will be acces
173173
```shell
174174
$ curl https://llama31.example.com/v1/chat/completions \
175175
-H 'Content-Type: application/json' \
176-
-H 'Authorization: Bearer &lt;dstack token&gt;' \
176+
-H 'Authorization: Bearer &lt;user token&gt;' \
177177
-d '{
178178
"model": "meta-llama/Meta-Llama-3.1-8B-Instruct",
179179
"messages": [

‎mkdocs/docs/examples/accelerators/amd.md‎

Lines changed: 9 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -93,8 +93,15 @@ Here are examples of a [service](../../concepts/services.md) that deploy
9393
</div>
9494

9595
!!! info "Docker image"
96-
AMD deployments require specifying an image that already includes ROCm
97-
drivers. The SGLang and vLLM examples above use pinned ROCm images.
96+
AMD workloads require specifying an image with ROCm-compatible userspace and
97+
framework packages. The SGLang and vLLM examples above use pinned ROCm
98+
images.
99+
100+
If you already have a ROCm-compatible image, use it. Otherwise, choose an
101+
image for the framework you use from
102+
[ROCm Docker images](https://hub.docker.com/u/rocm), e.g. `rocm/sgl-dev`
103+
for SGLang, `rocm/vllm` for vLLM, or `rocm/pytorch` for PyTorch. For
104+
generic AMD dev environments or tasks, use `rocm/dev-ubuntu-24.04`.
98105

99106
To request multiple GPUs, specify the quantity after the GPU name, separated by a colon, e.g., `MI300X:4`.
100107

‎mkdocs/docs/examples/accelerators/tenstorrent.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -98,7 +98,7 @@ at `<dstack server URL>/proxy/services/<project name>/<run name>/`.
9898
```shell
9999
$ curl http://127.0.0.1:3000/proxy/services/main/tt-inference-server/v1/chat/completions \
100100
-X POST \
101-
-H 'Authorization: Bearer &lt;dstack token&gt;' \
101+
-H 'Authorization: Bearer &lt;user token&gt;' \
102102
-H 'Content-Type: application/json' \
103103
-d '{
104104
"model": "meta-llama/Llama-3.2-1B-Instruct",

‎mkdocs/docs/examples/inference/nim.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -71,7 +71,7 @@ If no gateway is created, the service endpoint will be available at `<dstack ser
7171
```shell
7272
$ curl http://127.0.0.1:3000/proxy/services/main/nemotron120/v1/chat/completions \
7373
-X POST \
74-
-H 'Authorization: Bearer &lt;dstack token&gt;' \
74+
-H 'Authorization: Bearer &lt;user token&gt;' \
7575
-H 'Content-Type: application/json' \
7676
-d '{
7777
"model": "nvidia/nemotron-3-super-120b-a12b",

‎mkdocs/docs/examples/inference/sglang.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -112,7 +112,7 @@ If no gateway is created, the service endpoint will be available at `<dstack ser
112112
```shell
113113
curl http://127.0.0.1:3000/proxy/services/main/qwen36/v1/chat/completions \
114114
-X POST \
115-
-H 'Authorization: Bearer &lt;dstack token&gt;' \
115+
-H 'Authorization: Bearer &lt;user token&gt;' \
116116
-H 'Content-Type: application/json' \
117117
-d '{
118118
"model": "Qwen/Qwen3.6-27B",

‎mkdocs/docs/examples/inference/trtllm.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -72,7 +72,7 @@ If no gateway is created, the service endpoint will be available at `<dstack ser
7272
```shell
7373
$ curl http://127.0.0.1:3000/proxy/services/main/qwen235/v1/chat/completions \
7474
-X POST \
75-
-H 'Authorization: Bearer &lt;dstack token&gt;' \
75+
-H 'Authorization: Bearer &lt;user token&gt;' \
7676
-H 'Content-Type: application/json' \
7777
-d '{
7878
"model": "nvidia/Qwen3-235B-A22B-FP8",

‎mkdocs/docs/examples/inference/vllm.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -106,7 +106,7 @@ If no gateway is created, the service endpoint will be available at `<dstack ser
106106
```shell
107107
curl http://127.0.0.1:3000/proxy/services/main/qwen36/v1/chat/completions \
108108
-X POST \
109-
-H 'Authorization: Bearer &lt;dstack token&gt;' \
109+
-H 'Authorization: Bearer &lt;user token&gt;' \
110110
-H 'Content-Type: application/json' \
111111
-d '{
112112
"model": "Qwen/Qwen3.6-27B",

‎mkdocs/docs/examples/models/deepseek-v4.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -83,7 +83,7 @@ If no gateway is created, the service endpoint will be available at
8383
```shell
8484
curl http://127.0.0.1:3000/proxy/services/main/deepseek-v4/v1/chat/completions \
8585
-X POST \
86-
-H 'Authorization: Bearer &lt;dstack token&gt;' \
86+
-H 'Authorization: Bearer &lt;user token&gt;' \
8787
-H 'Content-Type: application/json' \
8888
-d '{
8989
"model": "deepseek-ai/DeepSeek-V4-Pro",
@@ -114,7 +114,7 @@ top-level JSON fields.
114114
```shell
115115
curl http://127.0.0.1:3000/proxy/services/main/deepseek-v4/v1/chat/completions \
116116
-X POST \
117-
-H 'Authorization: Bearer &lt;dstack token&gt;' \
117+
-H 'Authorization: Bearer &lt;user token&gt;' \
118118
-H 'Content-Type: application/json' \
119119
-d '{
120120
"model": "deepseek-ai/DeepSeek-V4-Pro",

‎mkdocs/docs/examples/models/qwen36.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -110,7 +110,7 @@ If no gateway is created, the service endpoint will be available at
110110
```shell
111111
curl http://127.0.0.1:3000/proxy/services/main/qwen36/v1/chat/completions \
112112
-X POST \
113-
-H 'Authorization: Bearer &lt;dstack token&gt;' \
113+
-H 'Authorization: Bearer &lt;user token&gt;' \
114114
-H 'Content-Type: application/json' \
115115
-d '{
116116
"model": "Qwen/Qwen3.6-27B",
@@ -138,7 +138,7 @@ To disable thinking, pass `chat_template_kwargs` in the request body.
138138
```shell
139139
curl http://127.0.0.1:3000/proxy/services/main/qwen36/v1/chat/completions \
140140
-X POST \
141-
-H 'Authorization: Bearer &lt;dstack token&gt;' \
141+
-H 'Authorization: Bearer &lt;user token&gt;' \
142142
-H 'Content-Type: application/json' \
143143
-d '{
144144
"model": "Qwen/Qwen3.6-27B",

0 commit comments

Comments
 (0)