| last_edited | 2026-09-26 |
|---|---|
| title | Configuration |
| description | Configuration file reference, environment variables, and file locations. |
Default location:
| Platform | Path |
|---|---|
| macOS / Linux | ~/.msgvault/config.toml |
| Windows | C:\Users\<you>\.msgvault\config.toml |
Override the data directory with the MSGVAULT_HOME environment variable or the --home flag (see below).
For a first archive, add only the sections required by your source. Optional provider setup is covered in recommended configuration. The complete example below illustrates the available sections; it is not a required starting configuration.
Starting in v0.20.0, remote deletion remains permanently opt-in. The invoking
CLI may enable it durably with [deletion] remote_enabled = true or for one
command with MSGVAULT_ENABLE_REMOTE_DELETE=1. Both mechanisms are permanent;
there is no planned automatic removal of the guardrail.
Consent belongs to the invoking CLI. When a command uses a remote daemon, the
CLI forwards its effective consent for that operation; the remote daemon's own
[deletion] section is not server policy for a command invoked elsewhere.
Staging, listing, inspecting, and dry-running deletion batches remain ungated.
Use recommended configuration for a guided setup, then return here for exact keys and defaults. Each processing feature has a separate scope and consent contract.
The defaults in this reference apply when a key is absent. They differ from the
values setup providers writes after confirmation: embeddings and people sweeps
start disabled, while setup can enable them and add schedules. Existing explicit
values remain in effect. Provider keys alone do not enable processing.
- Profile automation: tracked people, sweep providers, budgets, and fact resolution.
- Conversation briefs: enrolled people and versioned summaries through the sweep provider.
- External enrichment: Exa/SixtyFour policy setup, exact consent, request limits, and the persistent suppression key.
- Document indexing: extraction, document vectors, and separate query consent.
- Vector search: text, person, and visual indexes.
People sweeps use one named protocol profile at a time. A profile records the exact endpoint, model, wire protocol, negotiated output mode, privacy posture, and source scope. Built-in OpenAI, OpenRouter, and Venice presets bind the protocol, endpoint, and authentication scheme; you still choose the model and privacy policy. Msgvault never changes the active profile or switches providers automatically.
[people.sweep]
enabled = true
provider = "glm"
[people.sweep.providers.glm]
protocol = "openai_chat"
endpoint = "https://api.z.ai/api/paas/v4"
model = "glm-5.3"
auth = "bearer"
credential = "env"
credential_env = "ZAI_API_KEY"
output_mode = "prompt_json"
token_limit_parameter = "max_tokens"
reasoning_effort = "max"
request_timeout = "1m"
retention_posture = "provider-declared"
training_posture = "provider-declared"
allowed_sources = ["conversation_text", "meeting_text"]
source_since = "2026-01-01"
allow_sensitive = trueallow_sensitive = true is required for real sweeps: any packet that
carries seed or context evidence is marked sensitive, so a profile set to
false fails on every real sweep. true permits sending that verbatim
archive text to the selected provider; false leaves only the synthetic
capability check, which sends no archive text.
Usable protocols are openai_chat, openai_responses,
anthropic_messages, and google_generate_content. A fifth protocol,
codex_app_server, is defined but release-gated and cannot run yet; see the
Codex app-server profiles section below. Onboarding negotiates and saves
native_json_schema, json_object, or prompt_json. OpenAI Chat profiles
also save either max_completion_tokens or max_tokens; the other protocols
use their defined token-limit field.
Other providers use explicit protocol profiles; OpenRouter, Venice, and OpenAI also have built-in presets:
| Example profile | Protocol | Typical profile choice |
|---|---|---|
| GLM 5.3 | openai_chat |
Z.AI API base, glm-5.3, often max_tokens |
| Kimi K3 | openai_chat or anthropic_messages |
Choose the exact API surface the account exposes |
| OpenRouter | openai_chat |
OpenRouter API base and one explicit routed model ID |
| Venice | openai_chat |
Venice API base and one explicit model ID |
| open-agent-api | openai_chat |
The gateway's loopback API base and exposed model ID |
| Gemini | google_generate_content |
Google API base and one Gemini model ID |
| Anthropic | anthropic_messages |
Anthropic API base and one Claude model ID |
| OpenAI Responses | openai_responses |
OpenAI API base and one Responses model ID |
Confirm current endpoints, model identifiers, privacy terms, and subscription rules with the selected operator before saving a profile. OpenRouter and Venice may route a request to another upstream operator, so the profile's retention and training declarations must cover that full path. Logged-in or subscription-backed endpoints, including local gateways, must be used within their provider terms.
Credentials are not stored in this TOML. credential = "stored" keeps a
profile-specific secret under the private tokens directory and is supported
on Linux and macOS only; credential = "env" stores only the selected
environment-variable name and works everywhere. Environment-variable names are
host-only settings: configure them through the CLI or TOML, not the Web UI.
On hosts without stored-key support, the Web UI hides profile enrollment and
key fields. Run msgvault person provider add
with --credential-env on the daemon host, then reload the Web settings to
check and select the profile.
credential = "none" is restricted to credentialless local or Codex paths.
Changing a credential value does not change the profile fingerprint, but
changing its source or reference does.
Only msgvault person provider add may contact models.dev, and only when a
transport field (--protocol, --endpoint, --model, --auth) is missing
or --accept-catalog-prices is set; it sends no archive data or provider
credential. --custom skips that catalog entirely; the required synthetic
check still contacts the endpoint selected in the profile. The catalog is never used
by scheduled or manual sweeps. A catalog suggestion also never chooses where a
credential is sent: onboarding pairs a credential only with an endpoint you
passed explicitly via --endpoint or with the first-party API hosts compiled
into msgvault, so a compromised catalog cannot redirect your key. A successful check does not grant consent:
msgvault person provider consent <name> --yes is a separate explicit step.
Live credential checks are optional developer or operator verification and are
never CI requirements.
Enable model-assisted profile maintenance only after configuring, checking, and consenting to a provider. A person must also be tracked before the sweep maintains their facts. Conversation briefs use this same provider and schedule, with separate enrollment and interval controls.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Run the scheduled people sweep with the selected provider. |
provider |
default |
Name of a table under [people.sweep.providers]. The initial profile has an OpenAI endpoint but no model; it is not a usable, consented provider. Setup can explicitly select openai, openrouter, or venice, or configure local ollama. |
schedule |
15 2 * * * |
Daily at 02:15 in the daemon's time zone. An omitted or empty value receives this default; use enabled = false to disable the sweep. |
work_batch_size |
25 |
Tracked people considered in one worker batch. |
historical_message_cap |
2000 |
Maximum archived messages considered when finding context for each profile field. |
context_per_target |
8 |
Maximum context items selected for each profile field. |
evidence_max_bytes |
131072 |
Byte limit for an evidence packet. |
evidence_max_items |
200 |
Item limit for an evidence packet. |
backstop_interval |
24h |
Interval before checking tracked people for changes missed by incremental work. |
Hosted people inference requires an explicit setup providers --provider <openai|openrouter|venice> --model <model> choice, a credential source, and
explicit retention, training, and sensitive-content decisions. An
OPENAI_API_KEY alone configures only eligible embedding lanes. With no OpenAI
key or explicit provider choice, setup can offer the configured loopback Ollama
chat model when --allow-sensitive is supplied. An enabled sweep is preserved;
--provider then fails with instructions to use person provider add and
person provider use. See setup flags.
Request and token limits apply to sweeps and briefs. Setup keeps these defaults; it does not obtain provider prices or set a monetary limit.
| Scope | Request key and default | Input-token key and default | Output-token key and default |
|---|---|---|---|
| One person | max_requests_per_person = 4 |
max_input_tokens_per_person = 200000 |
max_output_tokens_per_person = 16000 |
| One run | max_requests_per_run = 100 |
max_input_tokens_per_run = 1000000 |
max_output_tokens_per_run = 160000 |
| One day | max_requests_per_day = 500 |
max_input_tokens_per_day = 5000000 |
max_output_tokens_per_day = 800000 |
max_estimated_cost_microusd_per_run and max_estimated_cost_microusd_per_day
both default to 0, which disables those cost limits. To use either, supply
positive input_cost_microusd_per_million_tokens and
output_cost_microusd_per_million_tokens; both price assumptions also default
to 0. Values are integer millionths of a US dollar. These are local estimates
from the prices you provide, not a provider billing limit.
Control how often enrolled people receive a "Last time we talked" brief and
how much text each generation uses. Briefs use the selected sweep provider and
share its budgets. The profile must permit sensitive content and include
conversation_text in allowed_sources.
Only supported chat and text-message sources supply brief evidence; email, meeting transcripts, documents, and your own replies are excluded. See the brief guide for supported sources and enrollment instructions.
Generation is enabled here by default, but each person must be enrolled separately. These settings apply to everyone; there are no per-person overrides. Restart the daemon after changing them.
[people.sweep.brief]
enabled = true
min_interval = "168h"
pre_call_window = "72h"
max_items = 40
max_bytes = 65536
overlap_items = 8
max_output_tokens = 2048
max_rendered_runes = 560| Key | Default | Description |
|---|---|---|
enabled |
true |
Generate briefs for enrolled people when the sweep runs. false stops generation for everyone without removing enrollments. |
min_interval |
168h |
Minimum age of the current version before a scheduled run regenerates it. A rejected current version counts as no brief and is replaced on the next eligible run. msgvault person brief generate bypasses this. Must be positive. |
pre_call_window |
72h |
Regenerate this far ahead of a due contact cadence, so the brief is current before you reach out. Must be positive. |
max_items |
40 |
Maximum archive items admitted to one brief window. Must be positive. |
max_bytes |
65536 |
Maximum packet size for one brief window. Must be positive. |
overlap_items |
8 |
Items already covered by the previous brief's window that may be re-admitted so a continued thread is recognizable. Must not be negative or exceed max_items. |
max_output_tokens |
2048 |
Output cap for the brief call. Must be positive and must not exceed [people.sweep.budgets] max_output_tokens_per_person, because a brief is one more call against the same per-person ceiling. |
max_rendered_runes |
560 |
Maximum length of the rendered paragraph, in Unicode runes. Msgvault enforces it after the model answers by dropping structured items from the tail: uncertainties first, then follow-ups, then highlights, never the last-interaction sentence and never mid-sentence. Every dropped item is counted in the version's dropped_item_count, and the stored structure and evidence pointers are the trimmed ones. Must be at least 240, which is the interaction summary's own maximum length. |
An invalid value fails configuration validation with the offending key named, rather than being clamped.
The codex_app_server protocol is not usable in this release. Its transport
stays unavailable until the executable isolation gate releases a verified
build, and until then every Codex operation fails closed with
codex app-server isolation is not released. The profile shape is documented
here so the configuration is ready when the gate ships. Codex sign-in and
model routes return HTTP 503 before changing credentials or consent. The
terminal-only person provider enroll-codex
command creates a new profile through the daemon; host-side person provider login reauthenticates the selected existing Codex profile. Both remain gated.
The Linux launcher disables Codex's local execution environment and shell tools. The app server can manage its staged OAuth credential, but its command and filesystem interfaces cannot access it. Command-execution requests or events abort inference. Enabling a release still requires a real authenticated structured-inference check through this launcher.
codex_app_server profiles are also the one protocol person provider add
cannot create: generic onboarding negotiates HTTP capabilities through an
endpoint, while codex_app_server has no endpoint to negotiate against and
runs through an attested local Codex executable instead. The following is
a reference shape for that gated implementation, not a working setup procedure:
[people.sweep.providers.codex]
protocol = "codex_app_server"
model = "gpt-5.3-codex"
auth = "none"
credential = "none"
reasoning_effort = "medium"
retention_posture = "zero_retention"
training_posture = "no_training"
allowed_sources = ["conversation_text"]
source_since = "2026-01-01"
allow_sensitive = trueendpoint is not allowed, and auth and credential must both be set to
"none": the transport is the local Codex app server, authenticated by its
own ChatGPT login. Sensitivity policy is protocol-agnostic: real Codex sweeps
send the same seed and context packets, so allow_sensitive = true is
required here as well. Login, model discovery, checks, and consent cannot make
this profile usable while the release gate is closed. Use one of the available
HTTP protocols for current profile automation.
TOML treats backslashes inside double-quoted strings as escape characters. On Windows, this means native paths like "C:\Users\you\..." will cause a parse error.
Use one of these formats instead:
# Forward slashes (recommended)
client_secrets = "C:/Users/you/Downloads/client_secret.json"
# Single-quoted string (backslashes are literal)
client_secrets = 'C:\Users\you\Downloads\client_secret.json'msgvault setup providers writes recommended values for the [vector],
[attachments.documents], and [people.sweep] sections from the API keys in
your environment, and msgvault setup status reports every lane with its
provider, model, consent state, and next step. The values it chooses are
listed in Recommended Configuration.
| Key | Default | Description |
|---|---|---|
data_dir |
~/.msgvault |
Base directory for all data |
export_dir |
{data_dir}/exports |
Directory for attachment ZIPs, downloads, and files opened from the TUI |
database_url |
{data_dir}/msgvault.db |
SQLite database path or PostgreSQL DSN |
loose_attachments |
false |
Keep attachments as loose files and reject pack/repack commands instead of creating immutable packs |
Attachments and OAuth tokens are stored in subdirectories of data_dir (attachments/ and tokens/ respectively). These paths are not independently configurable.
Setting loose_attachments = true prevents new pack files but does not
convert existing packs. Stop the daemon and run msgvault unpack-attachments
once to materialize their contents as loose files. Backup restore also restores
attachments loose while this setting is enabled.
Hosted extraction and local full-text indexing for standalone document attachments. It is disabled by default. Enabling it does not grant consent or send data: an operator must generate an authenticated capability manifest and record consent for the exact effective policy before a build can upload a document.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Allow explicit document extraction commands |
provider |
mistral |
Pinned extraction provider |
region |
eu |
Pinned provider region and EU endpoint |
api_key_env |
MISTRAL_API_KEY |
Environment variable containing the provider key |
model |
mistral-ocr-4-0 |
Pinned OCR model |
retention_posture |
unknown |
Confirmed provider posture: standard or zdr |
training_posture |
unknown |
Confirmed provider posture: default-opt-out or opted-out |
max_file_bytes |
52428800 |
Maximum original document size (50 MiB) |
max_pages_per_document |
500 |
Maximum provider units for one document |
max_response_bytes |
67108864 |
Maximum provider response size (64 MiB) |
max_normalized_chars |
25000000 |
Maximum locally retained normalized characters |
max_spool_bytes |
536870912 |
Maximum private staging-directory usage (512 MiB) |
min_free_space_bytes |
1073741824 |
Free space preserved before staging (1 GiB) |
request_timeout |
5m |
Timeout for each provider request attempt |
max_retries |
3 |
Maximum transient retries |
max_pages_per_run |
10000 |
Conservative provider-unit budget for one run |
max_estimated_cost_usd_per_run |
50 |
Cost-planning ceiling for one run |
estimated_cost_usd_per_1000_units |
0 |
Operator-supplied current price assumption; zero disables cost calculation |
pricing_assumption_on |
— | Date for the price assumption, in YYYY-MM-DD form |
| Key | Default | Description |
|---|---|---|
enabled |
false |
Convert standalone text/csv attachments locally to PDF before the authorized PDF upload; disabled conversion leaves raw CSV outside the provider-authorized scope |
CSV conversion uses Docbank's default record, cell, cell byte, and PDF limits, tightened by the configured original file, response, and page ceilings. The generated PDF is transient. The archive keeps the original CSV hash and MIME type plus the conversion receipt and page, record, and cell provenance. The conversion declaration participates in the exact consent fingerprint only when enabled.
Provider uploads are manual-only: msgvault serve does not schedule document
extraction. Each documents build or documents resume receives its capability
manifest explicitly and displays its upload and cost preflight before requiring
--yes. When document indexing is enabled, the daemon's weekly reconciliation
and local derivative cleanup remain automatic and make no provider requests.
[attachments.documents.scope] accepts message_types; an empty list includes
all supported standalone attachment sources. The first release requires
[attachments.documents.index].lexical = true and store_chunk_text = true.
Hosted document embeddings are not enabled by this configuration.
Enable document vectors separately with
[attachments.documents.index.embeddings] enabled = true, an enabled text
embedding provider, and distinct consents for document text and query text.
They use [vector.embed.schedule] for automatic embedding of already extracted
document chunks. That schedule never extracts new attachments. Setup enables
this subtable when both document extraction and a supported text provider are
selected, but leaves the two consent steps to you.
See Document Attachment Indexing for the complete probe, consent, build, and recovery flow.
| Key | Default | Description |
|---|---|---|
client_secrets |
— | Path to Google OAuth client_secret.json for browser OAuth flows |
service_account_key |
— | Path to a Google service account key JSON for Workspace domain-wide delegation |
Named OAuth apps for Google Workspace organizations that require their own OAuth credentials. Each entry can define a separate browser OAuth client_secret.json, service account key, or both. Use --oauth-app <name> with add-account to bind an account to a named app.
| Key | Default | Description |
|---|---|---|
client_secrets |
— | Path to the org's client_secret.json |
service_account_key |
— | Path to the org's Google service account key JSON |
See OAuth Setup: Google Workspace Accounts for when and why you need named apps.
Discord's --oauth-app value is only a protected bot-token binding label. It
is not resolved from this section and does not require an [oauth.apps] entry.
export-messages uses the same daemon and database configuration as other
archive commands. It does not load provider credentials or make provider API
calls. The older export-discord compatibility command has the same read-only
provider behavior.
When service_account_key is configured, msgvault add-account <email> validates the delegated Gmail profile and registers the account without storing a per-user refresh token. The service account key file must be owner-only on Unix-like systems, for example chmod 600 /path/to/service-account.json.
Connect through the CardDAV account workflow so the daemon validates discovery before saving these settings.
| Key | Default | Description |
|---|---|---|
provider |
"" |
Empty for a password-based server, or google for Google Contacts |
oauth_app |
"" |
Named Google OAuth app; empty selects [oauth] |
base_url |
"" |
CardDAV discovery URL; Google setup supplies its canonical URL |
username |
"" |
Server username or Google account email |
schedule |
"" |
Cron schedule; empty disables scheduled sync |
enabled |
false |
Enable the configured connection |
trusted_origin |
"" |
Exact HTTPS origin approved for private access, including its port; a trailing / is accepted. Applies only when it matches the account URL's origin. |
trusted_addresses |
[] |
Private IP addresses to dial for trusted_origin, without DNS. Accepts 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 100.64.0.0/10, and fc00::/7; rejects duplicates, IPv6 zones, loopback, and link-local addresses. |
Set both trusted-destination keys together. See the private-server setup for an example, restart requirements, and behavior when the origin does not match.
Passwords and Google tokens stay in the configured token directory, outside
config.toml. See Google Contacts setup
for browser and terminal authorization.
Configuration for Microsoft 365 / Outlook.com OAuth and Microsoft Teams Graph
sync. Required only if you use add-o365, add-teams, or sync-teams.
| Key | Default | Description |
|---|---|---|
client_id |
— | Azure AD Application (client) ID (required) |
redirect_uri |
http://localhost:8089/callback/microsoft |
OAuth redirect URI registered in the Azure AD app |
tenant_id |
common |
Azure AD tenant ID; common allows both personal and org accounts |
See OAuth Setup: Microsoft 365 for app registration steps. Teams uses the same client_id but requests Microsoft Graph scopes and stores tokens under tokens/teams_<email>.json; Outlook/Hotmail IMAP OAuth uses tokens/microsoft_<email>.json.
Optional source-scoped Fastmail JMAP identity inventory. This does not replace IMAP ingestion credentials: add and sync the mailbox normally, then use the API token only to discover masked and send-as addresses that belong to that source.
| Key | Default | Description |
|---|---|---|
source_id |
— | Positive numeric archive source ID; mutually exclusive with account |
account |
— | Unambiguous source identifier or display name; mutually exclusive with source_id |
api_token |
(required) | Fastmail API token used for the JMAP identity inventory |
auto_confirm_identities |
false |
Refresh and apply strong provider identity evidence after successful mailbox syncs |
Exactly one source selector is required. Prefer source_id when two sources
share an identifier or display name. With automatic confirmation disabled,
msgvault identity discover --source-id <id> --provider fetches the inventory
for an explicit preview; add --apply only after reviewing it. See People,
Profiles, and Source Identities.
Provider-wide Discord import settings and optional message-container filters.
Register guilds and store their bot credential first with msgvault add-discord; tokens and binding labels do not belong in config.toml.
| Key | Default | Description |
|---|---|---|
max_media_bytes |
52428800 (50 MiB) |
Maximum size of one Discord attachment downloaded during sync or backfill |
max_media_mb |
— | Same cap in MiB; when set it takes precedence over max_media_bytes |
media |
true |
Download attachment bytes at all |
media_scope |
all |
Which conversations collect media: all, direct (direct and group chats only), or none |
media_max_participants |
20 |
Skip media from conversations with more participants than this; 0 disables the cap. See Media policy |
edit_rescan_window |
168h (seven days) |
Trailing per-channel/thread window refreshed for edits, deletions, and reaction summaries |
Use an exact guild ID for a per-guild filter block. The same block also takes
the per-account media overrides (media, max_media_mb) that the other chat
providers put under accounts_config:
[discord.guilds."123456789012345678"]
include = ["456789012345678901"]
exclude = ["567890123456789012"]
# media = false
# max_media_mb = 25An empty include means every accessible text or announcement channel, thread,
and forum post. Top-level channels match directly. A child inherits its
parent's state unless its own ID appears explicitly. An explicit child include
can override an excluded parent; an explicit child exclude can override an
included parent. exclude wins when the same ID is in both lists. See
Discord.
[beeper], [slack], [discord], and [teams] share one attachment policy
vocabulary. It decides which chat media is downloaded during sync and backfill;
message text is always archived.
| Key | Default | Description |
|---|---|---|
media |
true |
Download attachment bytes. false archives messages without their media and records a policy_scope skip marker |
media_scope |
all |
all collects from every conversation; direct collects only from direct and group chats (not channels, rooms, or guild channels); none collects nothing |
media_max_participants |
20 |
Skip media from conversations with more participants than this. Omitting the key applies the default; an explicit 0 removes the cap |
max_media_mb |
250 (Discord 50) |
Per-attachment size cap in MiB. Sized for long voice notes, screen recordings, and phone video from direct chats now that the participant cap keeps large-room volume out |
accounts_config |
— | Per-account overrides of media and max_media_mb, keyed by Beeper accountID, Slack team ID, or Teams account email. Discord uses [discord.guilds."<id>"] instead |
The participant cap exists because most attachment bytes in a real chat
archive come from large rooms whose forwarded videos nobody wants kept. Direct
chats and small groups keep their photos, voice notes, and files. A skipped
occurrence is recorded with a typed marker (participant_threshold,
policy_scope, account_policy, or size_cap) that distinguishes a
deliberate skip from a failed download, so the backfill-*-media commands do
not retry it unless the policy changes.
[beeper]
media_scope = "all"
media_max_participants = 20
max_media_mb = 250
# Keep everything from one account regardless of room size or size cap.
[beeper.accounts_config.signal]
media = true
max_media_mb = 500
# Never download from another account.
[beeper.accounts_config.telegram]
media = falsePolicy changes apply to future downloads. Media already stored under an
earlier policy stays until you run msgvault purge-excluded-media, which
removes attachment bytes the current policy would no longer collect.
Structured file logging. Disabled by default. Enable it to get persistent, machine-readable logs for troubleshooting. Every CLI invocation writes a unique run_id on every log line so you can trace a single run across shared daily log files.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Turn on persistent file logging. Setting dir also enables it implicitly. |
dir |
<data_dir>/logs |
Directory for log files |
level |
info |
Log level: debug, info, warn, error |
sql_trace |
false |
Log every SQL query at info level (verbose, for debugging) |
sql_slow_ms |
100 |
Threshold in ms above which SQL queries are logged at warn level. 0 uses the built-in default (100 ms). |
Log files are named msgvault-YYYY-MM-DD.log (UTC date), written as newline-delimited JSON. When a daily log exceeds 50 MiB it rotates to .log.1, .log.2, etc. (up to 5 rotated files).
When SQL logging is enabled, slow/error entries include query arguments and streaming query durations, which makes it easier to diagnose expensive reads without enabling full trace output.
Use msgvault logs to view and tail log files from the selected local or remote daemon. See CLI Reference: logs.
| Key | Default | Description |
|---|---|---|
rate_limit_qps |
5 |
Scales Gmail's local quota-unit refill rate: 5 allows 250 units/second; 3 allows 150. Gmail values above 5 are capped at 5. Also sets Microsoft Teams Graph requests/second, without that cap, so lowering it slows Teams imports too. Reduce it if Gmail reports quota errors; Google's project quotas can be lower than this local budget. |
archive_remote_images |
false |
Download remote email images during Gmail/IMAP sync and EML, EMLX, MBOX, and PST imports |
trusted_imap_sent_mailboxes |
{} |
Per-IMAP-account Sent-folder names (keyed by the ACCOUNT identifier from msgvault list-accounts) that enable edited-copy snapshot refresh for servers without advertised special-use roles |
Remote image archiving is off by default. Enabling it contacts sender-controlled servers and can activate tracking pixels or disclose the archive server's IP address. Restart the daemon after changing the setting. It applies to new ingestion; existing mail needs an explicit backfill.
See remote email images for the opt-in workflow, supported formats, download limits, and effect on attachment counts.
trusted_imap_sent_mailboxes names each IMAP account's Sent folder for
servers — such as some Exchange/Outlook accounts — whose (possibly localized)
Sent folder advertises no RFC 6154 \Sent special-use role. A survivor copy in
that unadvertised folder never replaces the archived snapshot by default:
the sync adopts the surviving location but keeps the previously archived
body, raw MIME, recipients, and attachments, because a sender can forge any
RFC822 Message-ID. Placement the server does advertise — an unambiguous
\Sent or \Drafts role — remains trusted automatically regardless of
this setting. Naming a mailbox here states that it is that account's Sent
folder and holds only mail the account itself authored, which re-enables
refreshing the archived snapshot from an edited copy found there.
The mapping is keyed by the exact source identifier — copy the ACCOUNT
value printed by msgvault list-accounts (for example
imaps://user@example.com@imap.example.com:993), not the email or display
name. Trust never crosses accounts: a same-named mailbox in another synced
account stays untrusted, and an account with no entry has no explicit trust.
[sync]
trusted_imap_sent_mailboxes = { "imaps://user@example.com@imap.example.com:993" = ["Gesendete Elemente"] }This is an explicit trust assumption, not evidence: filters or IMAP rules
that file received mail into the listed mailbox would let that mail replace
an archived snapshot under the same Message-ID. Advertised unambiguous
\Sent and \Drafts placement is trusted automatically, and a mailbox
that carries \All, \Junk, or \Trash roles, or INBOX, is never
trusted — not even when listed here explicitly. A configured name the server
itself advertises as \Drafts keeps its Drafts meaning: explicit
configuration cannot turn a Drafts folder into the account's Sent folder.
Settings for the Web UI and API server started by msgvault serve. The same HTTP server is used by remote CLI access and by the local background daemon for archive-access CLI commands. The api_key setting is also reused for inbound bearer authentication when msgvault mcp --http starts a separate Streamable HTTP listener; that listener's address comes from the --http flag. See Web UI & API Server for API endpoint documentation and MCP Server for MCP client setup, or fetch /openapi.json from a running server for the generated OpenAPI contract.
| Key | Default | Description |
|---|---|---|
api_port |
0 (auto-select) |
Port the server listens on; 0 picks an open port at startup and clients discover it automatically. Set a fixed port for remote/NAS deployments. |
bind_addr |
127.0.0.1 |
Bind address |
api_key |
— | API key for daemon/API authentication and bearer authentication on msgvault mcp --http |
agent_access |
false |
Enable restricted agent grants; requires api_key to be non-empty. Read at daemon startup only; a config.toml edit takes effect only after a restart. |
allow_insecure |
false |
Allow non-loopback binding without api_key |
cors_origins |
[] |
Allowed CORS origins |
cors_credentials |
false |
Allow credentials in CORS requests |
cors_max_age |
0 |
CORS preflight cache duration in seconds |
trusted_proxies |
[] |
IP addresses or CIDRs allowed to supply forwarded HTTPS/host headers |
daemon_idle_timeout |
20m |
Idle timeout for lifecycle-managed background daemons; set to "0s" to disable |
daemon_auto_restart |
newer |
Local daemon restart policy when the CLI finds a different daemon binary version: newer, never, or always |
daemon_auto_start |
true |
Let CLI, TUI, and MCP commands start a local background daemon when none is running; set false when a supervisor runs msgvault serve |
daemon_idle_timeout applies only to background daemons started by msgvault daemon start or auto-started by a CLI command. Foreground msgvault serve keeps running until stopped. MSGVAULT_DAEMON_IDLE_TIMEOUT overrides the configured value for lifecycle-managed background daemons.
daemon_auto_restart = "newer" replaces an older compatible local daemon with the current CLI binary. Use "never" when another supervisor owns the daemon lifecycle, or "always" to restart whenever the recorded daemon version differs. Remote servers are never auto-restarted by a CLI client.
daemon_auto_start = false is for installs where a supervisor such as launchd, systemd, or Docker runs msgvault serve. Local archive commands then use the daemon that is already running, wait for one that is still starting, and otherwise fail with an error instead of starting their own. They also never replace a running daemon, whatever daemon_auto_restart says, because the supervisor owns restarts. msgvault daemon start, msgvault daemon restart, and the restart after msgvault update still start a daemon when you run them. Commands routed to [remote].url are unaffected.
Browser sessions are additive to API-key authentication. Existing CLI and
programmatic clients continue to send the configured key. For remote browser
access, terminate TLS at a reverse proxy and list that proxy—not arbitrary
clients—in trusted_proxies. See Web UI for the complete security
model and the plain-HTTP warning.
For MCP Streamable HTTP, send [server].api_key as Authorization: Bearer <key> on every /mcp request. This inbound credential is independent of
[remote].api_key, which authenticates msgvault mcp when it connects to a
remote daemon.
Defaults for the daemon-served browser application. These values can also be
changed from Settings; config.toml remains authoritative.
| Key | Default | Description |
|---|---|---|
default_search_mode |
full_text |
Initial mode: full_text, semantic, or hybrid |
theme |
system |
Color theme: system, light, or dark |
density |
compact |
Table density: compact or comfortable |
Browser-managed settings are validated and written with optimistic concurrency.
Only the [web] keys apply right away; every other config.toml category takes
effect after the daemon restarts, and the Settings page says so once per
category. Two things saved from the Settings page are not config.toml rows and
apply right away: person-enrichment provider API keys, and the CardDAV account,
which has its own save action. A "Saved. Restart the daemon" banner means the
file is saved but the running daemon still has its old value. Changing server.api_key requires a confirmation and takes
effect only after restart, which also invalidates browser sessions.
Optional provider-neutral integration for message-to-task links. Person agendas use the separate Kata connection below.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Enable discovery and capability checks |
endpoint |
— | Explicit loopback HTTP, Unix socket, or HTTPS endpoint; empty requests secure local discovery |
api_key |
— | Server-side credential; the browser sees only a masked hint of it, never the key |
default_project |
msgvault |
Fixed project used for create/link/search operations |
Remote plaintext HTTP is rejected. An endpoint is usable only when it supports the required idempotency and compare-and-swap capabilities; the UI distinguishes disabled, authentication required, incompatible, partial, stale, unavailable, and ready states.
Optional live person agendas backed by Kata. Tasks stay in Kata; msgvault shows their current state when you open a person's agenda. This integration is built against Kata v0.18.0 and requires Kata API schema version 0.21.0 or later.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Enable Kata person agendas |
endpoint |
— | Required when enabled: an explicit HTTPS URL, loopback HTTP URL, or Unix socket URL |
api_key |
— | Bearer credential sent by the daemon to Kata; Settings returns only its configured state and a masked hint |
default_project |
msgvault |
Existing active Kata project used for person agendas |
Create the project in Kata, then configure its endpoint and credential on the machine running the msgvault daemon:
[integrations.kata]
enabled = true
endpoint = "https://kata.example.com"
api_key = "replace-with-your-kata-api-key"
default_project = "msgvault"Restart the msgvault daemon after saving. These settings are also editable in
Settings and take effect after restart. Changing the endpoint to a different
origin in Settings clears its saved key unless you provide a replacement key
in the same save. Remote plaintext HTTP is rejected. Kata does not use local
endpoint discovery, and [integrations.tasks] does not configure person agendas.
Each Kata task can belong to one person. The scalar metadata value
msgvault.person is the person's canonical vCard UID, not their numeric
msgvault person ID. msgvault.list names its list and defaults to agenda.
Reads and unlink operations also recognize the person's UID aliases after a
merge.
Agendas show open tasks only, with at most 100 returned items. A truncated
response means more remain; open Kata to see them. The transport also limits
each response to 1 MiB and reports an error when it exceeds that limit.
Create, link, move between lists, and unlink tasks through msgvault; edit task
content or priority, complete tasks, and reopen them in Kata. See
person agenda commands.
Settings for daemon-side aggregate query behavior. The Web UI, TUI, MCP server, and aggregate list commands use these settings through the local daemon or a configured remote server.
| Key | Default | Description |
|---|---|---|
engine |
auto |
Aggregate engine: auto starts with live SQL and switches to DuckDB after cache maintenance succeeds; sql always uses live SQL; duckdb requires a usable Parquet cache |
auto_build_cache |
true |
Refresh a stale or missing Parquet cache automatically at startup, after scheduled or manual syncs, and when a query finds it due; false skips automatic builds. An explicit query --fresh or sync --build-cache can still request one |
min_rebuild_interval |
0s |
Minimum age of a usable cache before a sync, query, or daemon restart may queue an automatic rebuild. Queries serve the committed snapshot during the interval |
builder_memory_limit |
2GB |
DuckDB memory limit for cache builds, such as 4GB or 512MiB |
builder_threads |
min(CPUs, 2) | DuckDB threads for cache builds; zero keeps the default |
builder_temp_limit |
32GB |
Maximum spill-to-disk size for cache builds |
query_memory_limit |
512MB |
DuckDB memory limit for daemon aggregate queries; raise it on a large archive |
query_threads |
min(CPUs, 4) | DuckDB threads for daemon aggregate queries; zero keeps the default |
query_temp_limit |
2GB |
Maximum spill-to-disk size for daemon aggregate queries; a query that spills past it fails with a DuckDB out-of-memory error |
If a Web UI query runs out of memory or temporary disk space, its error names
the query-limit settings above. Try filters that narrow the results. On the
machine running msgvault, check available memory and free disk space before
raising these limits in config.toml. Restart the daemon to apply the change,
then retry the query. The limits cap resource use; they do not reserve memory
or disk space. Cache builds have separate builder_* limits.
The daemon starts HTTP health and API routing before analytics cache
maintenance. With engine = "duckdb", analytics remain unavailable until a
usable cache is ready. If no usable cache can be built or opened, msgvault serve
fails instead of silently falling back. A failed automatic refresh keeps serving
the last usable publication. With auto_build_cache = false, use
msgvault build-cache, query --fresh, or sync --build-cache for explicit
cache maintenance. Deprecated in 0.17.0:
per-command analytics flags such as msgvault tui --force-sql,
msgvault mcp --force-sql, msgvault tui --no-cache-build, and
--no-sqlite-scanner were replaced by this daemon-level section. Use
engine = "sql" to force live SQL.
min_rebuild_interval limits automatic refreshes requested by syncs, queries,
and daemon startup. A restart or sync within the interval leaves the existing
publication in service and schedules a rebuild check when the interval ends.
A busy archive can therefore serve Parquet analytics that lag SQLite by
approximately the configured interval plus cache build time. Deleted messages
can remain visible in analytics query results until the next cache publication,
including while the interval has not elapsed and while a build runs. A zero
interval still allows stale rows during the build. Use query --fresh to wait
for analytics that include deletions committed before the request.
Explicit refresh requests and recovery of an absent, interrupted, incompatible, or otherwise unusable cache are not delayed.
The daemon runs automatic rebuilds in the background, outside the scheduled
sync that requested them, so syncs keep their cadence while a build runs. A
sync that finishes during a build does not discard it: the build publishes
what it exported and marks the cache so the next build is a full rebuild.
This partial snapshot remains usable across daemon restarts, and its next
automatic rebuild still honors min_rebuild_interval.
Cache build memory and temporary disk usage scale with archive size, so a
minimum interval can prevent repeated archive-scale work when sources sync
frequently. Changes under [analytics] take effect after the daemon restarts.
This setting governs the aggregate views (Senders/Domains/Labels/Time) and is ignored entirely when [data].database_url points at PostgreSQL — a PostgreSQL backend always uses live SQL for those views, and build-cache refuses to run against it. It does not affect the Web UI's Explore, Files, or People/domains workspaces, which require the SQLite + DuckDB/Parquet cache regardless of this setting and are unavailable on PostgreSQL; see PostgreSQL Backend for the current scope.
Default settings for msgvault backup. See Backup for the
capture, verify, and restore workflow.
| Key | Default | Description |
|---|---|---|
repo |
— | Default backup repository directory used when a backup subcommand omits --repo |
zstd_level |
0 |
Compression level for backup pack files. 0 uses msgvault's built-in default; otherwise use 1 through 19 |
When set, archive-access CLI commands use the remote server by default. Without [remote].url, they use the local background daemon instead. Pass --local to use the local daemon instead of the configured remote.
| Key | Default | Description |
|---|---|---|
url |
— | Remote API base URL (e.g. http://nas-ip:8080) |
api_key |
— | API key used by remote commands |
allow_insecure |
false |
Allow HTTP remote connections |
Affected CLI commands include search (FTS mode), query, show-message, stats, list-accounts, list-senders, list-domains, list-labels, identity subcommands, collection subcommands, export-eml, export-attachment, export-attachments, and tui.
Scheduled sync sources for the web server. Each [[accounts]] entry defines a
cron schedule for automatic background syncing. Gmail, IMAP, Microsoft Teams,
and Discord sources are supported. For IMAP or Teams, use the account display
name/email when available rather than a raw provider identifier. Discord
schedules must use the exact guild ID because guild display names are mutable
and may be duplicated.
| Key | Default | Description |
|---|---|---|
email |
(required) | Account identifier/display name, or exact Discord guild ID, to sync |
schedule |
— | Cron expression for sync schedule (e.g., 0 * * * *) |
enabled |
true |
Whether scheduled sync is active for this account |
For example, schedule one previously registered Discord guild independently:
[[accounts]]
email = "123456789012345678"
schedule = "*/30 * * * *"
enabled = trueScheduled SMS Backup & Restore sources are configured with [[synctech_sms.sources]] entries. These are created automatically by msgvault add-synctech-sms-drive, but can also be edited directly.
| Key | Default | Description |
|---|---|---|
name |
(required) | Source name used by sync-synctech-sms <name> and scheduler logs |
enabled |
true |
Whether the source is active |
backend |
local |
local for a path on disk, or drive for Google Drive |
path |
— | Local XML/ZIP file or directory when backend = "local" |
folder_id |
— | Google Drive folder ID when backend = "drive" |
google_account |
— | Google account used for Drive access |
owner_phone |
(required) | Owner phone number in E.164 format |
schedule |
— | Cron expression used by msgvault serve |
include_sms |
true |
Import SMS records |
include_mms |
true |
Import MMS records |
include_calls |
true |
Import call logs |
include_attachments |
true |
Import MMS attachments |
stable_after |
10m |
How long Drive files must remain unchanged before import |
oauth_app |
— | Named Google OAuth app to use |
Scheduled Google Calendar sync is configured with top-level [[gcal]] entries. Each entry is one OAuth account; msgvault serve runs it on the given cron schedule (first run full-syncs and registers calendars, later runs are incremental). Authorize the account first with msgvault add-calendar.
[[gcal]]
name = "primary" # optional; defaults to email
email = "you@gmail.com" # OAuth account = token key
oauth_app = "" # optional named OAuth app
calendars = [] # optional calendarId filter; empty = owner+writer
schedule = "0 */6 * * *" # 5-field cron, no seconds
enabled = true| Key | Default | Description |
|---|---|---|
name |
Source name used by sync-calendar <name> and scheduler logs |
|
email |
(required) | Google account that owns the token (the token key) |
oauth_app |
— | Named Google OAuth app to use |
calendars |
— | Specific calendar IDs to sync; empty syncs owned/writable calendars |
schedule |
— | Cron expression used by msgvault serve |
enabled |
false |
Whether the source is daemon-scheduled |
Archive chats from a locally running Beeper Desktop. A single
block (not a list): the Beeper Desktop API is loopback-only, so there is one
instance per machine and the daemon must run beside it. Authorize first with
msgvault add-beeper.
[beeper]
# url = "http://localhost:23373" # Beeper Desktop API (default)
enabled = true # gate for the daemon schedule
schedule = "*/30 * * * *" # 5-field cron; empty = manual sync only
accounts = [] # accountID include filter (empty = all)
exclude_accounts = [] # skip networks archived natively, e.g. ["whatsapp"]
rate_limit_qps = 20 # request rate against the local API
media = true # download attachment bytes
media_scope = "all" # all, direct, or none
media_max_participants = 20 # skip media from larger rooms; 0 = no cap
max_media_mb = 250 # per-attachment download cap (MiB)
# [beeper.accounts_config.signal] # per-account override, keyed by accountID
# media = true
# max_media_mb = 500| Key | Default | Description |
|---|---|---|
url |
http://localhost:23373 |
Beeper Desktop API base URL |
enabled |
false |
Whether the daemon schedules Beeper sync |
schedule |
— | Cron expression used by msgvault serve |
accounts |
all | Beeper accountIDs to sync (include filter) |
exclude_accounts |
— | Beeper accountIDs to skip (wins over accounts) |
rate_limit_qps |
20 |
Request rate limit against the local API |
media |
true |
Download attachment bytes (failed downloads retry via backfill-beeper-media) |
media_scope |
all |
all, direct, or none; see Media policy |
media_max_participants |
20 |
Skip media from conversations above this many participants; 0 = no cap |
max_media_mb |
250 |
Per-attachment download cap in MiB (over-cap media is recorded as a size_cap skip and retried only after the cap changes) |
accounts_config |
— | Per-accountID media and max_media_mb overrides |
The daemon can send stored WAV and MP3 audio from any captured source, including messaging and email imports, to a separately running Docbank media service. The service needs Docbank's media HTTP routes. See Send audio to Docbank for the capture and processing rules.
[integrations.docbank]
enabled = true
url = "http://127.0.0.1:8080" # your Docbank daemon; the port is an example
api_key_env = "DOCBANK_API_KEY" # daemon environment variable with the key
all_sources_upload_consent = true # allow audio from every captured source to leave msgvault
# asr_profile = "asr" # optional Docbank profile for audio without source text| Key | Default | Description |
|---|---|---|
enabled |
false |
Schedule the stored-media job in msgvault serve |
url |
— | Docbank base URL: HTTPS, or HTTP on a loopback address. User info, query strings and fragments are rejected |
api_key_env |
— | Name of the daemon environment variable that holds the Docbank API key. It is read for each request and sent as X-Api-Key |
all_sources_upload_consent |
false |
Allow stored audio and explicit source transcripts from every captured source, including future providers, to be sent to url. Without it the job only records local state |
asr_profile |
— | Optional Docbank processing profile for stored audio without usable source text. An empty value retains audio without requesting processing. Msgvault rejects supplied-transcript, which Docbank reserves for supplied transcript input. |
The former Beeper-only upload_consent setting no longer enables uploads.
Existing users must explicitly set all_sources_upload_consent = true to resume
sending audio, including Beeper recordings.
The daemon reads these settings at startup, so restart it after a change. A new
url starts a separate delivery record; earlier rows stay. Disabling the route
stops the job and keeps its rows. A failed setup, such as an invalid url,
does the same and logs a warning. all_sources_upload_consent covers transport only; the
Docbank daemon's processing consent still decides whether a configured profile
may run. The route inspects stored CAS bytes, so MIME claims do not expand
Docbank's WAV and MP3 capability. Capture gaps and unsupported formats remain
typed local states.
Archive Slack workspaces. A single block covers every
registered workspace (tokens are per-workspace files). Authorize each
workspace first with msgvault add-slack.
[slack]
enabled = true # gate for the daemon schedule
schedule = "*/30 * * * *" # 5-field cron; empty = manual sync only
channels = [] # channel-name include filter (empty = all memberships)
exclude_channels = [] # channel names to skip, e.g. ["noise"]
dms = true # sync one-to-one direct messages
group_dms = true # sync group direct messages
media = true # download shared-file bytes
media_scope = "all" # all, direct, or none
media_max_participants = 20 # skip files from larger channels; 0 = no cap
max_media_mb = 250 # per-file download cap (MiB)
# [slack.accounts_config.T0123456] # per-workspace override, keyed by team ID
# media = false| Key | Default | Description |
|---|---|---|
enabled |
false |
Whether the daemon schedules Slack sync |
schedule |
— | Cron expression used by msgvault serve |
channels |
all | Channel names to sync (include filter; never applies to DMs or group DMs) |
exclude_channels |
— | Channel names to skip (wins over channels) |
dms |
true |
Sync one-to-one DMs; false pauses them without removing archived messages |
group_dms |
true |
Sync group DMs; false pauses them without removing archived messages |
media |
true |
Download shared-file bytes (failed downloads retry via backfill-slack-media) |
media_scope |
all |
all, direct (DMs and group DMs only), or none; see Media policy |
media_max_participants |
20 |
Skip files from conversations above this many members; 0 = no cap |
max_media_mb |
250 |
Per-file download cap in MiB (over-cap files are recorded as a size_cap skip and retried only after the cap changes) |
accounts_config |
— | Per-team-ID media and max_media_mb overrides |
Media policy for Microsoft Teams chats and channels. Teams
sync itself is scheduled through [[accounts]]; this table only decides which
attachments are downloaded.
[teams]
media = true
media_scope = "all"
media_max_participants = 20
max_media_mb = 250
[teams.accounts_config."user@example.com"]
media = true
max_media_mb = 500| Key | Default | Description |
|---|---|---|
media |
true |
Download attachment and inline hosted-content bytes (failed downloads retry via backfill-teams-media) |
media_scope |
all |
all, direct (chats only, not channels), or none; see Media policy |
media_max_participants |
20 |
Skip media from chats and channels above this many members; 0 = no cap |
max_media_mb |
250 |
Per-attachment download cap in MiB |
accounts_config |
— | Per-account overrides of media and max_media_mb, keyed by the Teams account email |
Granola meeting-notes sync is configured with top-level [[granola]] entries.
Each entry is one Granola account. identifier is a stable source label;
account_email is the primary identity used for organizer attribution.
msgvault serve runs it on the given cron schedule. Register the account
first with msgvault add-granola. See
Meeting Transcripts.
[[granola]]
identifier = "work" # stable label; defaults to "default" for a single entry
account_email = "you@example.com" # required primary account identity
api_key = "grn_..." # from the desktop app's settings (Business plan)
schedule = "0 */6 * * *" # 5-field cron, no seconds
enabled = true| Key | Default | Description |
|---|---|---|
identifier |
default (single entry) |
Source name used by sync-granola <identifier> and scheduler logs |
account_email |
(required) | Normalized primary account identity used for is_from_me |
api_key |
(required) | Granola API key (grn_…) |
schedule |
— | Cron expression used by msgvault serve |
enabled |
false |
Whether the source is daemon-scheduled |
Config loading preserves identifier and rejects labels without an effective
email, instructing you to add account_email. msgvault add-granola confirms
the primary identity even if aliases already exist. Manage aliases with
msgvault identity add <identifier> <email>, then run
msgvault sync-granola <identifier> --full after identity changes to repair
existing meeting attribution. A scheduled source must still be registered in
the archive; removing it prevents the scheduler from silently recreating it.
Circleback meeting sync is configured with top-level [[circleback]]
entries. Authentication is browser OAuth (msgvault add-circleback); no
secret lives in the config file. See
Meeting Transcripts.
[[circleback]]
identifier = "work" # stable label/token key; defaults to "default" for one entry
account_email = "you@example.com" # required primary account identity
schedule = "30 */6 * * *" # 5-field cron, no seconds
enabled = true| Key | Default | Description |
|---|---|---|
identifier |
default (single entry) |
Source name used by sync-circleback <identifier>, the token filename, and scheduler logs |
account_email |
(required) | Normalized primary account identity used for is_from_me |
endpoint |
production | MCP endpoint override (testing only) |
schedule |
— | Cron expression used by msgvault serve |
enabled |
false |
Whether the source is daemon-scheduled |
Config loading and alias repair follow the same rules as Granola: preserve the
label, add account_email, manage aliases with msgvault identity, and run
msgvault sync-circleback <identifier> --full after identity changes.
Circleback OAuth always confirms the primary identity; there is no identity
opt-out flag.
Notion meeting sync uses one top-level [[notion_meetings]] entry per Notion
identity. The token must belong to a read-only integration with AI Meeting
Notes access and Read Content access. User Information access is optional; it
is required only to resolve attendee IDs to verified email addresses.
[[notion_meetings]]
identifier = "notion-personal" # stable source label; defaults to "default" for one entry
account_email = "you@example.com" # required primary account identity
token = "ntn_..." # Notion integration token; keep this file private
schedule = "15 */6 * * *" # optional 5-field cron, no seconds
enabled = true| Key | Default | Description |
|---|---|---|
identifier |
default (single entry) |
Source name used by commands and scheduler logs |
account_email |
(required) | Normalized primary identity for relationships; it is not assumed to be the meeting organizer |
token |
(required) | Read-only Notion integration token |
schedule |
— | Cron expression used by msgvault serve |
enabled |
false |
Whether the source is daemon-scheduled |
Run msgvault add-notion-meetings <identifier> to validate access and register
the source before enabling a schedule. Removing the source prevents the
scheduler from recreating it. See Meeting Transcripts for
the 50-result discovery limit, attendee visibility, transcript retries, and
stored data.
Muesli meeting sync uses one top-level [[muesli]] entry per Muesli database.
The daemon reads the database read-only on its own host, so msgvault must run
on the Mac where Muesli records. No credential is needed.
[[muesli]]
identifier = "mac" # stable source label; defaults to "default" for one entry
account_email = "you@example.com" # required; you, the person who records
db_path = "~/Library/Application Support/Muesli/muesli.db" # optional; this is the default
phone_country_code = "1" # optional; convert national-format Contacts phones
schedule = "*/30 * * * *" # optional 5-field cron, no seconds
enabled = true| Key | Default | Description |
|---|---|---|
identifier |
default (single entry) |
Source name used by sync-muesli <identifier> and scheduler logs |
account_email |
(required) | Normalized primary identity; attributed as the organizer of every meeting |
db_path |
~/Library/Application Support/Muesli/muesli.db |
Muesli database path; ~ expands, and a relative path resolves against the config directory when --config is used |
contacts |
true |
Resolve attendees through Apple Contacts; needs Full Disk Access for the daemon |
contacts_path |
~/Library/Application Support/AddressBook |
Apple Contacts data folder; expands like db_path |
phone_country_code |
— | Country calling code (1–3 digits, such as "1" or "44") used for Contacts phone numbers typed without one; unset means only international numbers are used |
schedule |
— | Cron expression used by msgvault serve |
enabled |
false |
Whether the source is daemon-scheduled |
Run msgvault add-muesli <identifier> to check the database and register the
source before enabling a schedule. Run msgvault sync-muesli <identifier> --full
after identity changes to repair existing meeting attribution. Removing the
source prevents the scheduler from recreating it. See
Meeting Transcripts for what gets stored.
Top-level toggle and backend marker for semantic/hybrid search. SQLite vector search requires a build with sqlite_vec support (default via make build). PostgreSQL vector search requires a build with the pgvector tag and a PostgreSQL [data].database_url. See Vector Search for prerequisites, initial embedding, and the full workflow.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Turn on vector and hybrid search. When false, mode=vector and mode=hybrid return vector_not_enabled. |
backend |
sqlite-vec |
Backend marker. Supported values are sqlite-vec and pgvector; the concrete backend is selected from [data].database_url. |
db_path |
<data_dir>/vectors.db |
SQLite vector database path. Ignored by the PostgreSQL pgvector backend. |
skip_extension_create |
false |
PostgreSQL only. Skip CREATE EXTENSION IF NOT EXISTS vector when pgvector is already installed by an administrator. |
External OpenAI-compatible embedding endpoint used to convert message text into vectors. msgvault does not host a model; it calls the endpoint you configure. Use a local or self-hosted endpoint (Ollama, llama.cpp server, LM Studio, etc.) when message text must stay on your machine or network. Hosted endpoints also work but receive the text being embedded.
| Key | Default | Description |
|---|---|---|
api_format |
openai |
Request contract: openai (OpenAI-compatible /embeddings, one vector per message chunk) or voyage-contextual (Voyage /contextualizedembeddings; pins model = "voyage-context-4" and embeds chat conversation windows and turn-aware meeting chunks as contextual documents). |
endpoint |
(required) | HTTP(S) base URL for an OpenAI-compatible embeddings API. msgvault appends /embeddings (for example, set http://localhost:11434/v1, not .../embeddings). |
model |
(required) | Model name to pass in each request (e.g., nomic-embed-text). |
dimension |
(required) | Vector dimension. Must match the model's output dimension. |
document_prefix |
"" |
Model-specific instruction prepended to every document chunk after chunking (for example, "search_document: " for nomic-embed-text). The prefix does not reduce max_input_chars; maximum 4096 UTF-8 bytes. |
query_prefix |
"" |
Model-specific instruction prepended to every vector-search query (for example, "search_query: " for nomic-embed-text); maximum 4096 UTF-8 bytes. |
api_key_env |
— | Name of an environment variable containing the API key. Omit for anonymous endpoints. |
batch_size |
32 |
Embedding inputs per HTTP call. Long messages can contribute multiple chunk inputs. |
timeout |
30s |
Per-request timeout. |
max_retries |
3 |
Retries per batch on transient failures. |
max_input_chars |
32768 |
Character cap per embedding chunk, counted in characters rather than tokens. Too high and chunks are rejected or silently truncated; too low and long messages split into more chunks, adding embedding overhead. For example, start around 6000 for a 2k-token model such as Ollama's nomic-embed-text, then check representative content. See Matching max_input_chars to your embedder's context window. |
eta_window |
10 |
Number of recent progress samples used for ETA smoothing. |
Instead of naming an environment variable in api_key_env, you can store a
provider API key through Settings in the Web UI or the TUI. Stored keys live in
tokens/provider-credentials.json under the data directory with owner-only
file permissions. They are never written to config.toml. After saving,
Settings shows only a masked hint of the key, its first three and last three
characters, and whether it comes from the store or from the environment; the
key itself is never returned.
A stored key takes precedence over the environment variable named by
api_key_env. Each stored key is bound to the endpoint origin (scheme, host,
and port) it was saved for. If you later change the endpoint to another origin,
the stored key is removed automatically and must be entered again, so a key
is never sent to a host it was not entered for.
Changing a stored key for vector or multimodal (visual) embeddings requires a daemon
restart, like the other [vector] settings. Person enrichment and sweep keys
apply on the next run.
The index generation fingerprint includes the model, dimension, document and query prefixes, preprocessing settings, max_input_chars, embedding policy, and scope. Changing those settings triggers a stale-index error on the next vector/hybrid query. For an existing account-scoped generation built with CLI flags, set matching [vector.embed.scope].accounts and restart the daemon; otherwise run msgvault embeddings build --full-rebuild.
Controls text normalization before embedding.
| Key | Default | Description |
|---|---|---|
strip_quotes |
true |
Drop quoted reply blocks (> ... lines, reply preambles) before embedding. |
strip_signatures |
true |
Drop trailing signature blocks (content after -- ). |
strip_html |
true |
Convert HTML-only bodies to text and remove HTML markup before embedding. |
strip_base64 |
true |
Remove base64/data blobs before HTML stripping so encoded data does not crowd out prose. |
strip_url_tracking |
true |
Remove common tracking parameters such as utm_*, fbclid, and gclid from URLs. |
collapse_whitespace |
true |
Normalize repeated horizontal whitespace and blank lines. |
Hybrid ranking parameters applied at query time.
| Key | Default | Description |
|---|---|---|
rrf_k |
60 |
Reciprocal Rank Fusion constant. Higher values flatten score differences between signals. |
k_per_signal |
100 |
Candidate pool size drawn from each signal (BM25 or vector) before fusion. |
subject_boost |
2.0 |
Multiplier applied when a query term matches a message's subject line. |
max_page_size_hybrid |
50 |
Hard cap on page_size for vector/hybrid responses. Set to 0 to disable clamping. |
sqlite_accelerator |
auto |
Use a ready SQLite approximate index. Set to exact to keep exhaustive vector search. PostgreSQL ignores this setting. |
ann_nprobe |
8 |
SQLite index partitions searched per query. Higher values trade latency for recall. |
ann_oversample |
8 |
Approximate candidates requested per result before exact reranking. Range: 1–128. |
ann_threads |
CPU count, max 128 |
Native worker threads used by msgvault embeddings optimize. Range: 1–128. |
Accelerator tuning does not change the embedding generation fingerprint. It changes how stored vectors are searched or optimized, not how text is sent to the embedding provider.
Optional scope for newly built embedding generations. The zero value embeds the
full archive. A scoped generation embeds only matching messages.message_type
values:
[vector.embed.scope]
message_types = ["teams"]Scoped generations are intentionally partial. Vector and hybrid queries against
a scoped index must include a compatible message_type filter, such as
msgvault search "release planning" --mode hybrid --message-type teams; an
unscoped vector/hybrid query returns index_scope_mismatch instead of using the
partial index as if it covered the full archive.
| Key | Default | Description |
|---|---|---|
message_types |
[] (all types) |
Embed only messages of these types. |
accounts |
[] (all accounts) |
Embed only these accounts' messages, by canonical account identifier (display names are rejected here — they are not stable identities for a privacy boundary). Resolved to source IDs at startup; an unknown identifier fails vector initialization (or the CLI command). The daemon's scheduled embeds honor this scope, so it also acts as a privacy boundary: unlisted accounts' text is never sent to the embedding endpoint. |
accounts and message_types compose (both filters apply). The CLI flags
--account/--collection on msgvault embeddings build/resume override
accounts for a single run. Either scope dimension is part of the generation
fingerprint: changing it requires msgvault embeddings build --full-rebuild,
and because the fingerprint records archive-local source IDs, re-adding an
account under a new source ID also requires a rebuild. Account-scoped indexes
do not gate search the way message-type scopes do: out-of-scope accounts
simply have no vector matches and rank on BM25 alone in hybrid mode.
Optional background scheduling for the embed worker inside msgvault serve.
Empty config disables scheduled embedding; you can still run
msgvault embeddings build by hand.
| Key | Default | Description |
|---|---|---|
cron |
— | 5-field cron expression. Empty string disables the standalone cron. |
run_after_sync |
false |
Run an embed pass after successful scheduled Gmail, IMAP, Teams, and Discord syncs. Other sources use the standalone cron or a manual build. |
msgvault setup providers supplies run_after_sync = true and
cron = "*/15 * * * *" when it enables a text lane. It preserves either key
when explicitly set, including false and "". Already enabled text lanes keep
their schedules. See
Recommended Configuration.
Semantic people search: one curated document per durable person, built from
searchable non-sensitive attributes, embedded into the text-search generation.
Requires [vector] enabled = true and a separate consent
(msgvault person provider consent --semantic-embeddings --yes).
| Key | Default | Description |
|---|---|---|
enabled |
false |
Embed curated person documents and serve msgvault person search. |
retention_posture |
— | Your assertion about the embedding provider's retention; must be explicit (not unknown). |
training_posture |
— | Your assertion about the embedding provider's training use; must be explicit. |
Independently consented visual attachment lane over Voyage. Every value has a
default except the probe manifest; uploads fail closed without it, and a
daemon started with enabled = true and no manifest refuses every vector
lane until one exists.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Turn on the visual lane. |
provider |
voyage |
Only legal value. |
endpoint |
https://api.voyageai.com/v1 |
Pinned provider root; other origins are refused. |
api_key_env |
VOYAGE_API_KEY |
Environment variable holding the key. A key alone enables nothing. |
model |
voyage-multimodal-3.5 |
Pinned model. |
dimension |
1024 |
Pinned dimension. |
capabilities_file |
— | Manifest written by msgvault multimodal probe --seeds <dir> --out <file> --yes. |
max_context_chars |
4000 |
Owning-message text sent with each attachment. |
include_images |
true |
Embed still images (JPEG, PNG, WebP). |
include_animated_gifs |
false |
Embed animated GIFs; requires include_images and a manifest that authorized them. |
include_video |
true |
Embed direct-input MP4 video. |
allow_image_queries |
true |
Allow multimodal search --image. |
[vector.multimodal.scope] accepts the same message_types and accounts
keys as [vector.embed.scope]; [vector.multimodal.schedule] accepts the
same cron and run_after_sync keys as [vector.embed.schedule]. Consent is
recorded per generation by msgvault multimodal build --yes.
Dated activity projection and per-person contact state (first and last
contact, inbound/outbound, interaction count, inferred channel). It is the
deterministic source of "when did we last talk" for every person and runs
hourly by default inside msgvault serve. msgvault activity build runs it
by hand; --backstop rescans the whole archive.
| Key | Default | Description |
|---|---|---|
schedule |
17 * * * * |
5-field cron used by msgvault serve. Empty disables the scheduled job. |
timezone |
UTC |
IANA zone name for day bucketing. Local is rejected because the projection keys replay on the persisted zone name. |
max_direct_counterparts |
25 |
Largest conversation still projected as direct activity between its participants. |
batch_size |
500 |
Messages per projection batch. |
By default, msgvault stores everything under ~/.msgvault (macOS/Linux) or C:\Users\<you>\.msgvault (Windows). To use a different location, you have two options:
--home flag (per-command):
msgvault sync --home /mnt/data/msgvaultMSGVAULT_HOME environment variable (persistent):
export MSGVAULT_HOME=/mnt/data/msgvaultBoth options are equivalent: config.toml is loaded from the specified directory, and all data (database, tokens, attachments) is stored there. The --home flag takes priority over MSGVAULT_HOME.
The home or [data].data_dir directory may be a symlink to an existing
directory. Local daemon bookkeeping resolves the symlink and applies its
ownership and permission checks to the target directory.
| Variable | Description |
|---|---|
MSGVAULT_HOME |
Base directory for all data (default: ~/.msgvault) |
MSGVAULT_REMOTE_URL |
Remote URL for export-token (flag > env > config) |
MSGVAULT_REMOTE_API_KEY |
Remote API key for export-token (flag > env > config) |
All data lives under the msgvault home directory (~/.msgvault on macOS/Linux, C:\Users\<you>\.msgvault on Windows). The directory is created automatically on first use.
| File | Description |
|---|---|
config.toml |
Configuration file |
msgvault.db |
SQLite database (system of record when PostgreSQL is not configured) |
attachments/ |
Content-addressed attachment files |
tokens/ |
OAuth tokens per account |
logs/ |
Structured log files (when file logging is enabled) |
analytics/ |
Parquet cache files for Web UI and TUI analytical views |
Copy only the sections you need and replace example paths and credentials.
[data]
# Base data directory (default: ~/.msgvault)
data_dir = "/path/to/msgvault/data"
# User-requested exports (default: {data_dir}/exports)
export_dir = "/path/to/msgvault/exports"
# Database URL (default: {data_dir}/msgvault.db; PostgreSQL DSN supported)
database_url = "/path/to/msgvault.db"
# Keep attachment content as individual files instead of creating packs.
# loose_attachments = true
[oauth]
# Path to Google OAuth client secrets JSON for browser OAuth
client_secrets = "/path/to/client_secret.json"
# Google service account key for Workspace domain-wide delegation (optional)
# service_account_key = "/path/to/service-account.json"
# Named OAuth apps for Google Workspace orgs (optional)
[oauth.apps.acme]
client_secrets = "/path/to/acme_workspace_secret.json"
# service_account_key = "/path/to/acme_service_account.json"
[microsoft]
# Azure AD app registration client ID (required for M365)
client_id = "your-azure-app-client-id"
# redirect_uri = "http://localhost:8089/callback/microsoft" # default
# tenant_id = "your-tenant-id" # optional, default "common"
# Optional source-scoped Fastmail alias inventory.
[[fastmail]]
source_id = 14
api_token = "replace-with-a-Fastmail-API-token"
auto_confirm_identities = false
[discord]
# Per-attachment download cap (default: 50 MiB)
max_media_bytes = 52428800
# Skip attachments from rooms with more than this many participants
# (default: 20; 0 = no cap). Shared by [beeper], [slack], and [teams].
media_max_participants = 20
# Trailing edit/delete/reaction repair window (default: seven days)
edit_rescan_window = "168h"
[discord.guilds."123456789012345678"]
# Channel, thread, and forum-post IDs; empty include means all accessible.
include = ["456789012345678901"]
exclude = ["567890123456789012"]
[log]
# Persistent structured file logging (opt-in)
enabled = true
# dir = "/path/to/logs" # default: <data_dir>/logs
# level = "info" # debug, info, warn, error
# sql_trace = false # log every SQL query (verbose)
# sql_slow_ms = 100 # slow query threshold in ms
[sync]
# Gmail API rate limit (requests per second)
rate_limit_qps = 5
[server]
# API server settings (used by `msgvault serve` and `msgvault daemon`)
# api_port is optional; omit it (or set 0) to auto-select an open port that
# clients discover automatically. Set a fixed port for remote/NAS deployments.
api_port = 0
bind_addr = "127.0.0.1"
api_key = "your-secret-key"
daemon_idle_timeout = "20m" # background daemon idle timeout; "0s" disables
daemon_auto_restart = "newer" # newer, never, or always
daemon_auto_start = true # false when a supervisor runs msgvault serve
[analytics]
# Daemon-side analytics engine for Web UI, TUI, and aggregate HTTP views:
# "auto" starts on live SQL and switches to DuckDB after cache maintenance.
# "sql" always uses live SQL. "duckdb" requires a usable Parquet cache.
engine = "auto"
# Build a stale/missing cache during daemon startup and after scheduled syncs.
auto_build_cache = true
# Minimum age of a usable cache before a scheduled sync may rebuild it again.
# min_rebuild_interval = "6h"
[backup]
# Default repository for `msgvault backup`.
repo = "~/Backups/msgvault"
zstd_level = 0
[deletion]
# Durable consent for remote deletion execution. Opt in deliberately;
# defaults to false.
remote_enabled = false
[remote]
# Remote msgvault endpoint for CLI remote mode
url = "http://nas-ip:8080"
api_key = "remote-api-key"
allow_insecure = true
# Scheduled sync accounts
[[accounts]]
email = "you@gmail.com"
schedule = "0 * * * *"
enabled = true
[vector]
# Semantic and hybrid search (opt-in)
enabled = true
backend = "sqlite-vec"
# backend = "pgvector" # with a PostgreSQL database_url and pgvector build
[vector.embeddings]
endpoint = "http://localhost:11434/v1"
model = "nomic-embed-text"
dimension = 768
document_prefix = "search_document: "
query_prefix = "search_query: "
eta_window = 10
[vector.preprocess]
strip_quotes = true
strip_signatures = true
strip_html = true
strip_base64 = true
strip_url_tracking = true
collapse_whitespace = true
[vector.embed.scope]
# Empty means embed the full archive. Set this for partial generations.
message_types = ["sms", "mms"]
# Use stable account identifiers, not numeric source IDs. This keeps a scoped
# generation usable after a daemon restart.
# accounts = ["you@work.example"]
[attachments.documents]
# Hosted extraction is opt-in and requires a separately recorded consent.
enabled = false
provider = "mistral"
region = "eu"
api_key_env = "MISTRAL_API_KEY"
model = "mistral-ocr-4-0"
retention_posture = "zdr"
training_posture = "opted-out"
max_file_bytes = 52428800
max_pages_per_document = 500
max_response_bytes = 67108864
max_normalized_chars = 25000000
max_spool_bytes = 536870912
min_free_space_bytes = 1073741824
request_timeout = "5m"
max_retries = 3
max_pages_per_run = 10000
max_estimated_cost_usd_per_run = 50
# Set both pricing fields together to include a cost estimate in manual build preflight.
# estimated_cost_usd_per_1000_units = 0.001
# pricing_assumption_on = "2026-08-17"
[attachments.documents.scope]
# Empty includes every supported message type.
message_types = ["email"]
[attachments.documents.index]
lexical = true
store_chunk_text = true
[[synctech_sms.sources]]
name = "phone-backups"
enabled = true
backend = "drive"
folder_id = "google-drive-folder-id"
google_account = "you@gmail.com"
owner_phone = "+14155551234"
schedule = "30 4 * * *"