A playback engine core for Apple platforms.
FFmpeg demuxes. Apple plays. Dolby Atmos and Dolby Vision survive the trip.
No view, no controls, no state — you keep your AVPlayer and your UI.
Any container libavformat can read (MKV, MPEG-TS, AVI, …) is remuxed on the fly
into HLS-fMP4, served from a loopback HTTP server on 127.0.0.1, and handed
to a plain AVPlayer as a playlist URL. Not a single frame is re-encoded. Apple's
stack then does what only it can do on its own platforms:
- hardware decode (VideoToolbox),
- Dolby Atmos passthrough — EAC3+JOC is stream-copied, never decoded to PCM,
- Dolby Vision display engagement and the HDMI handshake,
- Match Content (frame rate + dynamic range) on tvOS,
- system-side HDR10 / HLG tone-mapping,
- and everything that rides along for free because it is AVPlayer: Picture in Picture, AirPlay, the Now Playing surface, spatial audio.
On Apple platforms the usual choice is between AVPlayer — deep OS integration, but only Apple's containers — and an mpv- or VLC-derived engine, which plays almost anything but renders its own frames and decodes audio to PCM, so Dolby Vision and Atmos die at its door. PrismCore is the third option: FFmpeg's container breadth in front of Apple's own playback stack.
For the handful of sources whose video AVPlayer cannot decode at all (VP9,
MPEG-2, VC-1, interlaced H.264), there is a software path — libavcodec into
AVSampleBufferDisplayLayer — so a host has one engine to call rather than two.
- Aether — a native media player for Apple platforms. PrismCore is the engine every non-Apple container routes through: ahead of libmpv, behind plain AVFoundation for the files it can already open. The remux path is on by default there; the software path is opt-in.
Shipping something on PrismCore? Open an issue and it gets listed here.
| Area | Summary |
|---|---|
| Containers | Anything libavformat demuxes — MKV, MP4, MPEG-TS, AVI, WebM, FLV, … |
| Video (native) | H.264, HEVC (incl. Main 10), AV1 where the device has a hardware decoder — asked with VTIsHardwareDecodeSupported, not inferred from a chip name |
| Video (software) | VP9 / VP8, MPEG-2, MPEG-4 Part 2, VC-1 / WMV3, AV1 without hardware (libdav1d), interlaced H.264 with CPU bwdif deinterlace at field rate |
| HDR | HDR10 (PQ) and HLG, signaled honestly in the master playlist and tone-mapped by the system |
| Dolby Vision | Profile 5 (dvh1 sample entry), 8.1 and 8.4 (hvc1 + SUPPLEMENTAL-CODECS), Profile 7 converted to single-layer 8.1 through libdovi; 8.2's Rec.709 base plays as plain SDR because that is what it is |
| HDR10+ | ST 2094-40 SEI is stream-copied like any other SEI. Opt-in detection (hdr10Plus: .standard on SourceProbe.open) reads a bounded run of video packets and reports seen / notSeenWithinBudget / unknown in SourceInfo.hdr10Plus — never "absent". Reporting only: no playlist or display-criteria signalling until a device run shows it helps |
| Dolby Atmos | EAC3+JOC stream-copied, and the dec3 box's TS 103 420 type-A extension re-applied to the init segment — without it AVFoundation plays the same bitstream as plain DD+ |
| Audio (copy) | AAC, AC3, EAC3, FLAC, ALAC — bit-for-bit |
| Audio (bridge) | TrueHD / MLP / DTS / DTS-HD MA / MP3 / MP2 / Opus / Vorbis / PCM → EAC3 5.1, 128 kbps per channel. Needs an FFmpeg build with the eac3 encoder; without it those sources take the software path instead, which decodes them itself |
| Multi-audio | Every viable track becomes an HLS alternate rendition with its language, name and channel count, so AVPlayer gets a real AVMediaSelectionGroup to switch on. preferredAudioLanguage: decides which one is DEFAULT, so playback starts in the right language instead of switching visibly after it. The software path switches too: SoftwarePlaybackPipeline.selectAudioTrack(streamIndex:) swaps the decoder mid-playback without touching the clock or the picture |
| Dialogue boost | Opt-in (dialogueBoost: on the session): extra "Dialogue Boost" renditions derived from the default track — which is the preferredAudioLanguage: track when one matched — decoded, centre channel favoured (bed −6 dB / −12 dB), re-encoded to EAC3 — marked public.accessibility.enhances-speech-intelligibility so hosts find them by characteristic. Engine-side because AVFoundation ignores audioMix/audio taps on HLS items. The base track stays bit-for-bit (Atmos included). Needs the eac3 encoder and a centre-channel source; stereo would need dialoguenhance, which current builds don't ship |
| Subtitles (text) | SubRip / ASS / SSA / WebVTT / mov_text converted during the remux read into segmented WebVTT renditions, cut on the video's own boundaries — so text survives PiP and AirPlay instead of living in a host overlay. ASS inline italics / bold / underline become WebVTT tags; \an / \pos placement and a WebVTT track's own cue settings ride the timing line, so a caption authored at the top of the frame stays there. External .srt / .vtt register as first-class renditions |
| Subtitles (bitmap) | PGS / DVB / DVD read by on-device Vision OCR into the same rendition machinery. Lossy by design — typography dies, text survives — and the raw tracks stay surfaced for a host that wants to draw them pixel-accurately |
| Closed captions (CEA-608) | Captions embedded in the video — A/53 cc_data in H.264 / HEVC SEI, the form US broadcast recordings, MPEG-TS captures and many disc rips carry — decoded during the remux read into the same WebVTT rendition machinery, so CC1…CC4 appear as real AVMediaSelectionOptions and survive PiP, AirPlay and external display. Pop-on, roll-up and paint-on, the full basic / special / extended character sets, labelled by channel and by the video track's language where it declares one. Caption bytes are reordered from decode to presentation order before decoding, which is what keeps them from drifting on any source with B-frames. A bounded packet scan before the first segment decides whether a source has captions at all — one that does not pays nothing in the copy loop. Styling (colour, italics, underline) and cell-accurate positioning are dropped on purpose, so the system caption renderer applies the viewer's own accessibility style. The master says CLOSED-CAPTIONS=NONE, so AVPlayer does not also surface the raw in-band captions as a second track; inBandClosedCaptions: true declares the scouted services as TYPE=CLOSED-CAPTIONS instead |
| Closed captions (CEA-708) | Not decoded. DTVCC packets are recognised in cc_data and skipped. A 708 service decode means the window model — eight windows with their own anchors, sizes, pen states and row locks — and a partial one puts text on screen in the wrong place while presenting itself as a caption track. Because effectively every 708 encoder also emits the 608 compatibility bytes, this costs nothing on real content; a stream that carries only 708 gets no caption rendition rather than a broken one |
| Seek & cache | Keyframe-aligned segment plan published upfront, demand-driven production with re-anchoring, absolute-tfdt continuity across restarts, byte-budgeted retention (1 GiB default; an evicted segment is reproduced on demand, so the budget bounds disk, not seekability) |
| Chapters | Matroska Chapters / MP4 chapter tracks reported as SourceInfo.chapters and PrismCoreSession.chapters (title + start/end seconds) — HLS cannot carry them, so they are the host's to draw as timeline markers and skip controls |
| Display | tvOS HDMI handshake driven by the engine: preferredDisplayCriteria programmed and settled before the item is loaded, which is the only ordering tvOS accepts for HDR HLS |
| Scrub previews | SeekPreviewService decodes the keyframe covering any position into a CGImage for a custom player HUD — its own context, CPU-only, cached per keyframe, and independent of which engine is playing. The trick-play answer for sources with no server-generated previews |
| Streaming | HTTP headers ride the demux connection (a Plex token, a WebDAV authorization), reconnect on dropped connections. A planned session whose source fails mid-film re-opens it and carries on from the segment it was producing (up to three times a minute; see When the source fails mid-film) |
| Custom input | A host that holds the bytes itself — an SMB mount through its own client, a debrid/torrent session, an encrypted store, a file inside a disc image — implements PrismCoreInput (read / seek / length) and passes a factory as input: to the session, the probe or SeekPreviewService. One instance per open, so the probe, the producer and a scrub preview never share a cursor; host errors surface as PrismCoreInputError. An input with no length is refused (.notSeekable) rather than serving a plan it cannot honour. Omit it and the engine reads exactly as before |
.package(url: "https://github.com/Wenzlik/PrismCore.git", from: "3.0.0")One call probes the source, picks the path and hands back something already started:
import PrismCore
switch try await PrismCoreEngine.open(url: mkvURL) {
case .remux(let session, let playlist):
player.replaceCurrentItem(with: AVPlayerItem(url: playlist)) // AVPlayer plays it
…
await session.stop()
case .software(let pipeline):
hostView.layer.addSublayer(pipeline.displayLayer!) // we play it
pipeline.play()
…
pipeline.stop()
}The remux path is the one that matters — it is what buys hardware decode, Atmos passthrough and Dolby Vision. Drive it directly when you want the detail:
let session = try PrismCoreSession(url: mkvURL, display: .current())
// optional, before start(): a sidecar joins the WebVTT renditions
try await session.addExternalSubtitle(url: srtURL, language: "cs", name: "Čeština")
let playlistURL = try await session.start() // .../master.m3u8 or .../index.m3u8
// hand playlistURL to your AVPlayer
…
await session.stop()The URL is a master playlist when the source has audio (that is where the selectable renditions live) and a media playlist when it hasn't. Treat it as opaque: the shape is a property of the source, not of the API.
start() can take seconds on a slow origin. Register before it and you get
the stages as they land, each with the time since the call:
let session = try PrismCoreSession(url: mkvURL, display: .current())
let checkpoints = try await session.startupCheckpoints() // before start()
Task {
for await mark in checkpoints {
switch mark.phase {
case .sourceOpened: status = "Opening…"
case .streamInfoResolved(let info): status = info.video?.codecName ?? "…"
case .segmentPlanReady(let origin, _): status = origin == .sequential
? "Indexing on first play…" : "Preparing…"
case .firstVideoSegmentWritten: status = "Starting playback…"
case .playlistServable: break
}
log("\(mark.elapsed) \(mark.phase)") // where the twenty seconds went
}
// The stream ends here — on success, on failure, and on stop().
}
let playlistURL = try await session.start()There is no percentage, on purpose: nothing can know in advance how long a probe over a slow origin takes, and this engine does not report numbers it cannot measure. Stages with timestamps are things that happened.
PrismCoreEngine.decide(for:) is exposed separately, so a host can ask which
path a source would take — and unit-test its own routing — without standing up
either engine.
A host that knows which language the viewer wants says so when it builds the session, and the served master starts in it:
let session = try PrismCoreSession(
url: mkvURL,
display: .current(),
preferredAudioLanguage: "cs", // "cze", "ces", "cs-CZ" mean the same
preferredSubtitleLanguage: "cs"
)Without them the rendition that carries DEFAULT is whichever track the
source listed first, so a viewer who wants Czech mounts the item, hears
English, and switches — a visible wrong-language moment at every start, and on
the remux path a switch also costs a rendition fetch.
- Audio. The matching track becomes the master's
DEFAULTrendition, above every other signal the engine uses to guess (the container's original and default flags, the demuxer's "best stream"). Because dialogue boost derives from the default track, it derives from this one. - Subtitles. The matching rendition is the only one ever marked
DEFAULT=YES,AUTOSELECT=YES, which is what makes AVKit engage it at load rather than starting with subtitles off. A full rendition wins the flag over a forced one of the same language;FORCEDitself is untouched. Don't pass this together withsetTimedTextCueHandlerunless the host suppresses its own overlay — otherwise AVKit and the host both draw the cues. - Matching is tolerant. ISO 639-2/B (
cze), 639-2/T (ces) and 639-1 (cs) are one language; a bare tag matches a regioned one (pt↔pt-BR) and an exact region wins over a bare one;undand an empty tag are not languages and match nothing. - A no-match is a no-op, never an error and never an empty selection: the source's own default stands. Nothing is ever dropped — every viable track is still an alternate rendition — and no decode, bridge or stream-copy decision changes. A preferred track this build can neither copy nor bridge is passed over, because a rendition AVPlayer cannot play is worse than the wrong language.
setTimedTextCueHandler streams every embedded cue (text, OCR'd bitmap,
CEA-608) as the remux produces it, and a handler registered late — or again,
to pick up a new setSubtitleDelaySeconds — is first replayed everything
produced so far. That replay means the session keeps every cue, and its dedup
key, for its whole lifetime. For a film that is a few thousand short strings;
for a long captioned stream, or a disc with many OCR'd tracks, it grows with
the text. A host that only draws a running overlay can bound it:
let session = try PrismCoreSession(
url: mkvURL,
display: .current(),
subtitleCueHistory: .bounded(maxCues: 1_024, maxBytes: 1 << 20)
)
let history = session.subtitleCueHistoryStats // retained cues/bytes, evictions.completeis the default and changes nothing.- Live delivery is never bounded. Only the replay is: a late handler gets the newest retained cues, with the delay in force now.
- The byte bound is UTF-8 of each cue's text plus its dedup key (stream, source times, text), so a cue costs a little over twice its text. The oldest cue leaves together with its key.
- Dedup covers the retained window only. Seek back far enough that a region's cues were evicted and they are delivered again — a host opting in must tolerate a repeated cue.
- A single cue larger than
maxBytesis delivered live and never retained. Cues produced before the presentation origin is known wait in a queue bounded the same way; anything it drops was never delivered. - Nothing is dropped silently:
evictedCuescounts what is no longer replayable,droppedBeforeOriginwhat never went out. - Both limits are clamped to at least 1, and
Options.subtitleCueHistoryreads back the bound in force. The policy ridesmakeSession(changing:)and both rejection fallbacks. It bounds the cue tap's history, not every byte of subtitle memory: WebVTT renditions are produced as before, and the software path has its own cache (below).
SoftwarePlaybackPipeline publishes the probe's metadata on the live player.
Audio selection works while playing or paused; the decoder is opened before
replacing the current one, the audio renderer is flushed, and the clock and
video renderer stay in place. The bounded rewind uses the playhead and the
configured audio delay. A refused rewind joins at the demux read position;
network and decoder latency can leave an audible gap. No gapless claim is made.
let audioTracks = pipeline.selectableAudioTracks // language, title, channelCount
let subtitleTracks = pipeline.selectableSubtitleTracks // embedded text only
let audioSelection = pipeline.selectedAudioStreamIndex
let subtitleSelection = pipeline.selectedSubtitleStreamIndex // nil = Off
if let track = audioTracks.first {
pipeline.selectAudioTrack(streamIndex: track.streamIndex) { accepted in
// Completion is on the feed queue. Dispatch UI work to the main queue.
}
}
pipeline.selectSubtitleTrack(streamIndex: subtitleTracks.first?.streamIndex)
pipeline.selectSubtitleTrack(streamIndex: nil) // Off
// In the host's periodic UI update, replace the overlay with this snapshot.
// An empty array means clear it; overlapping cues can produce several entries.
let captions = pipeline.activeSubtitleCuesUse source stream indices, not positions in the menu array. Read the settled
selection after completion. A same-track selection succeeds without a flush;
invalid indices and calls outside paused/playing state are refused. An audio
open failure preserves the current track. A later read/decode failure reports
false and .failed; inspect failureError. Do not call synchronous stop()
or load() from feed-queue callbacks.
Text starts Off, including forced tracks: the host chooses its language policy.
All supported text packets use the existing HLS text converter and are retained
in a bounded look-ahead cache, so selecting an already-read caption (or Off)
does not seek, flush A/V, or start a paused clock. TimedTextCue timestamps here
use the source axis, matching pipeline.currentTime; remux cue callbacks
use the origin-rebased AVPlayer axis. Cue payloads can contain WebVTT inline tags
and entities; the host supplies rendering. TimedTextCue.placement carries the
placement the source asked for (TextCuePlacement: a numpad alignment, plus an
anchor point from \pos or a WebVTT line:/position: pair) — nil for the
host's default, which is what nearly every cue wants. Polling clears expired
cues even at EOF or while the demuxer is waiting for data.
The cache holds at most 1,024 cues / 1 MiB of UTF-8 text across all tracks; expired cues are removed and excess incoming cues are dropped. Cues need valid PTS and positive duration. Seeking clears the cache and repopulates from the landing keyframe: a long caption whose packet precedes that keyframe may be missing until the next cue. Bitmap/OCR and external subtitle selection remain outside this software surface. ASS colours, fonts, karaoke and drawing overrides are not preserved by the text converter — only the inline styles and placement WebVTT has words for.
Over HDMI the panel's mode has to be programmed before AVPlayer sees the
playlist. tvOS validates an HDR variant's VIDEO-RANGE against the panel's
current mode synchronously, so a PQ master handed to an SDR-parked panel fails
outright (-11848 / -11868) instead of switching or tone-mapping. AVKit's
appliesPreferredDisplayCriteriaAutomatically cannot help here: it derives
criteria from the chosen variant's format description, which only exists after
the variant passes the very validation the switch has to precede. So it must be
false, and the order is engine-driven:
let controller = DisplayCriteriaController(window: window) // tvOS only
let session = try PrismCoreSession(url: source, display: .current())
let playlist = try await session.start()
if let choice = await session.displayCriteria {
controller.apply(choice) // 1. program the panel
await controller.waitForSwitch() // 2. let the handshake settle
}
player.replaceCurrentItem(with: AVPlayerItem(url: playlist)) // 3. only then load
player.play()
// on teardown: controller.reset()session.displayCriteria is clamped to the display the session was built for: a
non-DV panel is asked for the base layer's range, a non-HDR panel for a rate-only
switch — which is also what makes Match Frame Rate engage on SDR content.
After the handshake, controller.currentPanelIsHDR() is the value to feed back
into DisplayCapabilities(panelIsCurrentlyHDR:) for the next session.
Built-in panels (iPhone, iPad, Mac) engage HDR on demand and skip all of this.
PrismCoreSession.isMasterRejection(_:) plus makeMuxedFallbackSession() stays
as the backstop for the one state no API can prove: Match Content switched off on
an HDR-capable panel.
PrismCoreError.classify(_:) turns anything this engine (or the host's
AVPlayer) threw into one machine-readable case, so a failure can be acted on
rather than logged:
do {
let playlist = try await session.start()
…
} catch {
switch PrismCoreError.classify(error) {
case .originRefused: await refreshToken()
case .originRateLimited(_, let after, _): await backOff(after ?? 5)
case .videoCodecNotRemuxable: routeToSoftwarePath()
case .masterRejectedByPlayer: try await session.makeMasterRejectionFallbackSession()
default: show(error)
}
}The cases: originRefused (401/403/407), originRateLimited (429/503/509, with
the origin's own Retry-After), originUnreachable, noVideoStream,
videoCodecNotRemuxable, videoCodecUnplayable, startupBudgetExpired,
masterRejectedByPlayer, workDirectoryOutOfSpace, ffmpeg (raw code and
message, for the libav* failures with no honest mapping) and unknown.
retryability is three-valued — .retryable, .permanent, .unknown — because
for a startup budget or a full volume this engine genuinely cannot say, and a
Bool would have to invent an answer. After startup, session.remuxFailure is
the same classification of session.remuxError.
Every case is something the engine observed. Where it cannot separate two
situations they share one case: an origin that never answered and one that
vanished mid-session are both originUnreachable, because the evidence at the
failure site is identical — the host knows which it was from whether start()
had returned.
For a bug report, await session.diagnosticReport() gathers the rest in one
Codable value: PrismCore and FFmpeg versions (with any ABI mismatch), the
platform, the source's description and container layout, the routing verdict
and its reason, the session's options, every startup checkpoint with its time
and the plan origin (planned or sequential), the probe's phase timings, and
the run so far: the newest 64 playback events, timestamp repairs, audio
rendition production, the object-audio and Dolby Vision conversion findings,
retention evictions and the classified failure. The session keeps the
checkpoints and events itself, so the report has them even if the host never
subscribed. A source that never got a session (declined, or sent to the
software path) gets the same report from its probe:
PrismCoreEngine.diagnosticReport(for: probed, decision: Result { try PrismCoreEngine.decide(for: probed.info) }). jsonData() writes it with
sorted keys, so two reports diff line by line. schemaVersion names the shape,
and fields are only ever added.
What a report leaves out. It is built to be pasted into a public issue, so it
never carries HTTP headers (only how many there were), the URL's query string,
fragment or user:password@, the LAN access token, or any path on the
device's disk: the work directory, keyframeIndexCacheDirectory (reported as
on or off) and sidecar subtitle files are all left out, and a local source is
named by its file name alone. Error descriptions are also scrubbed of every one
of those values, because an error can quote the failing URL — and any URL it
quotes, a redirect target included, loses its credentials, query and fragment
too. Two things do stay
in: a remote URL's host and path, which are what a reproduction needs, and the
track titles, languages and file name the media itself carries. A host whose
URLs keep a token in the path should cut it before sharing.
A read that fails after start() — a Wi-Fi roam, a host range proxy
restarting, a 5xx from a media server, a share that blinked — no longer has to
end the session. In a planned session (one with a keyframe plan, which is
what makes seeking demand-driven) the producer closes the source, waits, opens
it again exactly the way startup does (avformat_find_stream_info included),
checks that the streams are the same ones (index, codec, time base), and
re-anchors at the segment it was producing. The partial fragment is dropped and
that segment is produced again from its keyframe, as for a seek; the timeline,
the init segment and every rendition — a lazy one already armed included —
carry on. Fetches that arrive meanwhile are held the way a seek holds them, so
AVPlayer sees a slow segment, not an error.
- What is retried: failures whose
retryabilityis.retryable(429/503/ 509, a 5xx, a dropped connection) and.unknown(how a host input's own error arrives)..permanentones (404, 403 and the other 4xx) are not: the same request gets the same answer. An.unknownfailure that recurs right after a successful re-open, before a segment has landed, is not retried either — the source is reachable and the bytes still do not read. - How often: after 1, 2 and 4 s, and at most three re-opens in any 60 s,
whatever the classification.
stop()ends a backoff or a re-open at once. - A demuxer that skips instead of failing is caught too. Matroska resyncs to the next cluster after a failed read and reports success with a hole behind it; a planned session reads the failure off the I/O context on every packet and repairs the hole instead of serving it.
- Sequential sessions (no plan: a live source, a Matroska whose index did not load) behave as before: the failure ends the remux and the playlist stays valid up to the last segment written.
- A host
PrismCoreInputfactory is called again for each re-open — seePrismCoreInputFactory: it must hand out a fresh instance over the same bytes, from any thread, at any time.
playbackEvents() reports both outcomes:
for await event in session.playbackEvents() {
switch event {
case .producerRecovered(let attempt, let cause): log("reconnected (\(attempt)/3) after \(cause)")
case .producerFailed(let failure): showError(failure) // == session.remuxFailure
default: break
}
}producerFailed is sent exactly once per session that fails — mid-film or at
startup, planned or sequential — and never for a stop(); remuxFailure is
already set when it arrives. prismcore-cli serve prints both.
┌─► video ──► fMP4 segments + index.m3u8
Source URL ──► libavformat demux ───┼─► audio 0 ─► audio0/{init.mp4,seg*.m4s,index.m3u8}
├─► audio N ─► audioN/… │ tmp dir
└─► subs N ─► subsN/seg*.vtt │
master.m3u8 ───────────────┤
▼
LoopbackHTTPServer ──► AVPlayer
The segmentation is ours. MPVKit's libavformat is built without the hls muxer,
so FMP4SegmentWriter drives the mp4 muxer with frag_custom and cuts where we
say (video keyframes at or after each 6 s boundary), and MediaPlaylistWriter
turns those fragments into a playlist. Two production modes exist:
- Planned (demand-driven) — the default for seekable VOD.
SegmentPlanmaps every segment upfront from the demuxer's keyframe index (trusted only past two witnesses: max keyframe gap 60 s, coverage ≥ one target duration), completePLAYLIST-TYPE:VODplaylists are published before the first packet, and a fetch outside the producer's window re-anchors it —av_seek_frameto the anchor keyframe, fresh muxers withfrag_discont+avoid_negative_ts=disabledsotfdtcarries absolute time and the produced segment sits exactly where the playlist promised. After EOF the producer parks and keeps answering demand for segments a re-anchor skipped. - Sequential — an EVENT playlist growing head-to-EOF (or, opted in, a sliding window — see below). Used when no trustworthy plan exists (live, unknown duration, junk index) and for the muxed-with-bridge shape, where re-anchoring would reset an encoder mid-fragment.
A source whose container carries no usable index (a Matroska without Cues, any
MPEG-TS) is stuck sequential on its first play — but the producer reads every
packet anyway, so with keyframeIndexCacheDirectory set on the session it
harvests the keyframe map as a by-product and persists it (bounded LRU). The
next play of the same source plans on that map from its first second, as if the
file had an index — and a cache hit skips the index-load seek entirely.
A map is only as good as the proof that the file has not changed under it, so
each one is bound to an identity: the full URL including its query (a query
can select the media, not only carry a token), the byte size, the container
duration, and the version the open actually read — a local file's mtime, or the
strong ETag a remote origin reported through the coordinated HTTP reader
(coordinatedHTTP), together with the URL that served it — after a redirect,
the target, since a tag is only unique per resource. Only a SHA-256 digest of that is written; the URL and any
token in it never reach the disk. The scrub preview (SeekPreviewService)
derives the same identity from its own open. A remote source that cannot prove
its version neither uses a stored map nor stores one, and plans from the source
every time: an origin or proxy with no ETag, only a weak W/ one, or only
Last-Modified; FFmpeg's own HTTP; and a host PrismCoreInput on a non-file:
URL. A file replaced on the server under the same URL (new strong ETag), or a
rotated query token, costs one rebuild rather than a wrong map. Sidecars written by
3.4.0 and earlier are ignored and rebuilt once.
A host that knows what it is playing can name it. The URL-based identity
cannot survive an address that changes while the bytes do not — a localhost
range proxy on a new port with a new token after every app launch is a miss
every time. keyframeIndexCacheIdentity keys the map on the host's own name
for the source instead: a namespace (the server), the item, and a strong
ETag the host has from its server's metadata.
let session = try PrismCoreSession(
url: proxyURL, // port and token change per launch
display: .current(),
keyframeIndexCacheDirectory: cachesDirectory,
keyframeIndexCacheIdentity: SourceCacheIdentity(
namespace: serverID, item: itemID, etag: item.etag
)
)
// The scrub preview must be given the same identity to share the map.
let preview = SeekPreviewService(
url: proxyURL, keyframeIndexCacheDirectory: cachesDirectory,
keyframeIndexCacheIdentity: SourceCacheIdentity(namespace: serverID, item: itemID, etag: item.etag)
)With it set, the URL and whatever the transport reports stop counting: it works
over FFmpeg's own HTTP and a host PrismCoreInput as well as coordinatedHTTP.
The key still includes the cache format version, the byte size and the
container duration. A weak (W/…) or empty ETag, or an empty namespace or
item, turns the cache off for that source — it does not fall back to the URL.
The host vouches for the tag, so it must change whenever the file does. A
session without it behaves exactly as before; a clone carries it.
A source that has an index can still play its first play sequentially:
SegmentPlan.indexLoadBudget (3 s) bounds the index-load seek, and behind a
range proxy one tail request for Matroska Cues can take longer than that on
its own. When the plan fell back because the budget ran out (not for a junk
or missing index), the session reads the index again in the background once its
first video segment is out — a second open with its own interrupt guard and a
60 s budget, cancelled by stop(), whose reader takes the origin only while
the producer has nothing in flight. If that open proves it read the same
version (the identity above is equal) and the index reaches the end of the
source, the map is stored as complete right away rather than at EOF, and the
session reports PlaybackEvent.segmentPlanAvailable(segments:). The running
session stays sequential; a successor (makeSession(changing:) with the same
cache directory, then a seek to the current time) plans as VOD from the map.
This runs only where the map can be stored, over coordinatedHTTP (with a
strong ETag or a host identity) and a keyframeIndexCacheDirectory, and only for a container
whose framing names where its index is (a Matroska SeekHead pointing at its
Cues); anywhere else the second read would be a scan of the whole file, and the
EOF harvest gets that map for free.
A sequential segment cannot be produced a second time — there is no plan to
re-anchor on — so the byte budget (segmentCacheBytes) deletes old files
behind the playhead and they are gone. By default the playlist does not say
so: it is EXT-X-PLAYLIST-TYPE:EVENT, which may never lose an entry, so it
keeps offering the deleted segments and a seek back to one is a 404 AVPlayer
counts as a failed segment. A host can ask for a window that tells the truth
instead:
let session = try PrismCoreSession(
url: sourceURL,
display: .current(),
sequentialPlaylist: .slidingWindow(seconds: 300) // default: .event
)It decides nothing about a planned session (its VOD playlist lists the whole source), and clones and both rejection fallbacks carry it. In the window:
- Shape. Every media playlist — the variant, each audio and subtitle
rendition — is written without
PLAYLIST-TYPEfrom its first version, and entries leave the front:EXT-X-MEDIA-SEQUENCEadvances by one per entry removed, segment names and media timestamps never change, andTARGETDURATIONstays the longest entry ever listed. There are no discontinuities in this output, so noEXT-X-DISCONTINUITY-SEQUENCE. - Size. At each cut the front moves to
secondsof content before the segment AVPlayer fetched last, and further if the budget evicted beyond that — never past that segment itself, never so far that a playlist lists less than three target durations, and not at all before the first fetch. - One instant for every playlist. All playlists drop the same stretch of
time. An audio rendition folds a boundary that carried no audio into its
next entry (and numbers its own files), so where its entry straddles the
cut the cut moves back to that entry's start, for everyone. Media sequence
numbers may therefore differ between renditions by those folds; the
playlist timeline is what lines them up. A rendition whose first audio
arrives after the window has moved starts on the window's front: its first
entry covers only the stretch still listed elsewhere, not the whole time
before it. That entry still sets its
TARGETDURATION, and so its three target durations of minimum window, for the rest of the session. A rendition that stops delivering (an audio track ending early) no longer holds the window back once its last entry falls behind the variant's own three-target floor: its entries leave with everyone else's and, if it resumes, it rejoins on the front the same way. - Unlink after a grace period. A segment leaves the playlists first. Its
files —
.m4sand.vttalike — are deleted only after its own duration plus the longest any of the session's playlists has been (RFC 8216 §6.2.2's rule, rounded up), so a client still reading an older manifest is not handed a 404. The period is checked at the next cut, not on a timer: while the producer is parked (a paused player) or finished, removed files wait, and whatever is still waiting at EOF goes withstop(). - The budget is soft. It counts what the playlists offer. Files in their
grace period are on disk on top of it, and the three-target-duration floor
can keep an evicted segment listed (and on disk) a little longer. A
playlist holds
secondsplus the up to 30 s the producer runs ahead, and the grace period is about as long again, so size the budget for roughly twicesecondsplus a minute at the source's bitrate; a smaller one wins oversecondsand the window shrinks toward the segment fetched last. - EOF closes whatever window is left with
EXT-X-ENDLIST.
A segment that has left the window and been deleted still answers 404 — what changed is that no current manifest offers it any more.
The loopback server speaks HTTP/1.1 with keep-alive (bounded per connection and by
an idle timeout), GET + HEAD, and pipelined requests. Payloads come from a
SegmentProvider rather than straight off disk, and a provider that answers
.pending gets an early 200 + Transfer-Encoding: chunked once a serve passes
2 s — keeping response headers inside AVPlayer's ~3.5 s media watchdog window. A
pending serve that ultimately fails aborts the connection (a truncated transfer
makes AVPlayer retry) instead of framing a cacheable empty 200.
The server binds 127.0.0.1 by default, which is right for playback on the
device and wrong for AirPlay: when the host routes to an Apple TV or an
AirPlay 2 TV, the receiver fetches the playlist and every segment itself,
and 127.0.0.1 resolves to the receiver. The master playlist — native WebVTT
renditions, alternate audio, the lot — is simply unreachable.
Opt in per session when, and only when, that is the route:
let session = try PrismCoreSession(
url: sourceURL,
display: .current(),
reachability: .localNetworkUnencryptedForAirPlay
)
let playlist = try await session.start()
// http://10.0.0.7:51234/<32-char token>/master.m3u8- Interface —
getifaddrs, IPv4 only, up and running, no loopback and no point-to-point links; tunnels (utun,ipsec,ppp), the peer-to-peer radios (awdl,llw,nan), Apple silicon's internalanpilinks and self-assigned169.254/16addresses are excluded outright. A realeninterface wins, then anything unrecognized, and an Internet Sharing or VMbridgelast; ties break on the interface's own number, so the choice is deterministic. The server binds that one address rather than0.0.0.0, so a VPN or a shared-internet bridge is never exposed. No interface left →start()throwsLoopbackHTTPServer.NoLocalNetworkInterfaceinstead of publishing a URL nobody can reach. - IPv6 is deliberately not supported. A literal needs brackets in a URL, a
link-local one needs a
%zonereceivers handle inconsistently, and the platform rotates temporary privacy addresses on its own schedule — which would make "the address changed mid-session" routine. Every AirPlay receiver on a home network is reachable over IPv4. - Token — 192 bits from the system CSPRNG, base64url, as the first path
component of every URL (
X-PrismCore-Tokenis accepted too, for clients that can set headers; an AirPlay receiver cannot). Everything without it is404, ahead of the method check, and a wrong token is indistinguishable from a wrong path so the server is not an oracle. The token buys the session's namespace and nothing else: path traversal, the method restriction, the request-line and header caps, the per-connection budget and the idle timeout all behave exactly as on loopback. - The address changing mid-session (Wi-Fi to Ethernet, a DHCP change, the
radio dropping) is watched with
NWPathMonitor. The server does not re-bind — the URL is already baked into theAVPlayerItemand every segment reference — it answers503and publishessession.serviceAddress == .addressLost(…). A host that sees that should stop the session and start a new one. An address that comes back resumes serving.
The residual risk, plainly: this is cleartext HTTP on the local network. Anyone on that LAN who can observe the traffic sees the token, the playlist and the media bytes, and anyone holding the token can fetch the session's segments for as long as it runs. The token makes the server unguessable, not private. Enable the mode for the duration of an AirPlay route on a network the user trusts, and stop the session when the route ends. Sessions that never AirPlay should never enable it — the default is unchanged and unreachable off-device.
The things that were expensive to learn, kept here so the next person doesn't pay for them twice.
Atmos needs a box FFmpeg drops. The CHANNELS="16/JOC" playlist attribute is
necessary but it is not what engages Atmos. AVFoundation takes the Dolby/MAT route
only when the dec3 box in the init segment carries the TS 103 420 type-A
extension — and FFmpeg's mp4 muxer drops that extension on a stream copy (plain
ffmpeg -c copy does the same). So EAC3Syncframe reads complexity_index_type_a
out of the bitstream's addbsi and EAC3Configuration.patch appends it to the
produced init segment, re-framing every enclosing box's size. Field placement
follows libavcodec's ac3_parser.c rather than a reading of the spec, and that
mattered three times over: the JOC signal can sit in a dependent substream
(Blu-ray-style DD+ puts it there, behind an AC-3 core frame), a dependent substream
carries a chanmap field an independent one doesn't, and both the converter-sync
bit and the whole mixing-substream block exist on independent substreams only. Each
omission shifted the walk by a bit or two — and a shifted walk does not fail, it
returns a confident wrong number.
hvc1 is a promise the record has to keep. A stream-copied HEVC track is given
an hvc1 sample entry explicitly, because FFmpeg's mp4 muxer defaults HEVC to
hev1 and Apple's HLS rules want hvc1. But hvc1 also asserts that every
parameter set lives in that entry, and Matroska CodecPrivate routinely says
otherwise (array_completeness = 0) — so HVCCNormalizer rewrites the record to
VPS/SPS/PPS in order with completeness asserted. It runs twice: once on the
input extradata, and again on the produced init segment, because movenc doesn't
copy our record into the sample entry, it rebuilds one and re-zeroes the flag while
still naming the entry hvc1. Sources that keep their parameter sets in band
(numOfArrays = 0, common in MP4 and MPEG-TS) get them harvested off the first
keyframes and filled in, since an empty hvcC under an hvc1 entry promises Apple
a decoder configuration that isn't there.
Dolby Vision is three separate claims that must agree. The manifest's CODECS,
the sample entry's fourcc, and the dvcC/dvvC box all describe the same track,
and AVPlayer checks them against each other. Profile 5 gets dvh1 as its sample
entry — its picture is IPT-PQc2, not YCbCr, and an hvc1 entry over P5 renders the
familiar green-and-purple — while 8.x deliberately keeps hvc1, because its base
layer really is plain-HEVC-compatible and the dvvC box is what upgrades it.
Claiming dvh1 there would deny the very fallback that makes 8.1 worth having.
The boxes themselves need -strict unofficial: they are Dolby's specification
rather than ISO's, movenc refuses to write them by default, and a DV source without
them is HDR10 with extra bytes.
WebVTT timing states a fact, not a convention. HLS bridges cue times to the
media timeline with X-TIMESTAMP-MAP=MPEGTS:<t>,LOCAL:<local>. Our fMP4 segments
carry the source's own stream-copied timestamps, so the media timeline starts at the
first video PTS — there is no MPEG-TS 10 s convention to honour, and assuming one is
exactly what makes fMP4 subtitles render ten seconds late in other implementations.
PrismCore writes cue times relative to the presentation origin and repeats
MPEGTS:round(origin × 90000),LOCAL:00:00:00.000 in every segment.
A flagged interlaced source is often lying. Broadcast H.264 is routinely flagged interlaced around progressive frames, so the probe decodes a handful of frames before believing the flag. Evicting those from the native path would trade hardware decode and Atmos passthrough for deinterlacing nothing.
Two read-only additions for hosts whose media comes from a server that has already analysed the same file. Both are opt-in and both default to today's behaviour exactly.
Exporting. SourceProbe.open(url:structure:) fills
ProbedSource.structure with the container's byte layout and seek index —
where the metadata region ends, where the first media element starts, whether
the container's index lives at the head or the tail, and (at .full) the
index itself with a verdict on whether it reaches the end of the file. The
export is for an out-of-process probe reading a local descriptor: .layout
re-reads the head and .full also pays an index-load seek, which is why
.none is the default and a routing probe over a network pays neither.
Nothing in it is inferred. Every field is optional or has an unknown case,
and IndexLocation.none / IndexCompleteness.absent need positive evidence
that a container declares no index — an empty index table at open is not that
evidence. A consumer across a network cannot tell a measurement from a
plausible guess, so this side does not make guesses.
let probed = try SourceProbe.open(url: url, structure: .full)
probed.structure.firstClusterOffset // Int64? — where media begins
probed.structure.indexLocation // .head / .tail / .none / .unknown
probed.structure.index?.completeness // .complete / .partial / .absent / .unknownHDR10+. No container declares it, so the only honest answer comes from the bitstream, and it costs reads — which is why it is asked for, like the structure export, and never paid by a routing probe by default.
let probed = try SourceProbe.open(url: url, hdr10Plus: .standard)
probed.info.hdr10Plus?.verdict // .seen / .notSeenWithinBudget / .unknown(reason)
probed.info.hdr10Plus?.applicationVersionnotSeenWithinBudget means exactly that. The finding changes nothing about
how the source is played: VIDEO-RANGE stays PQ, and HDR10+ rides the
stream-copy whether or not anyone asked.
Consuming. SourceProbe.open(_:hints:) takes what a caller already knows.
hints: nil is the unhinted open, unchanged.
let probed = try SourceProbe.open(url, hints: SourceOpenHints(
headerBytes: 4312, firstClusterOffset: 5184, indexLocation: .tail
))
probed.hints.firstReadBytes // what the first read was actually bounded to
probed.hints.rejections // why a hint was not used — never an errorA hint may make this engine do less work; it may never make it skip a check.
headerBytes and firstClusterOffset size the first read and nothing else, so
a stale value costs a read rather than a wrong parse. expectedValidator and
keyframes exist and are validated in full — the transport's ETag is checked
once before any byte is delivered, and a supplied map is checked against the
stream, its exact time base, monotonicity, a cap and the container's bounds —
but a surviving map is carried on ProbedSource.hints, not yet consumed by the
planner, and is never written into the local keyframe sidecar. A map
computed elsewhere and a map harvested by this machine's own read of the file
are the same numbers with different provenance, and the disk must not lose the
difference.
Deliberate omissions, so you don't have to read the source to find them:
- No UI. No view, no transport bar, no controls, no HUD.
- No published playback state. The remux path is a service: it hands you a URL
and gets out of the way. Your
AVPlayerremains the source of truth for time, rate and status. - No subtitle rendering. Text becomes WebVTT renditions AVPlayer selects itself; bitmap tracks are surfaced for a host that wants to draw them.
- No playlist or queue management. Make a new session for the next title.
- No analytics. Nothing leaves the device, and nothing phones anywhere.
- No re-encoding of video, ever. If a frame would have to be touched, that source belongs on the software path or somewhere else entirely.
| Min | |
|---|---|
| iOS | 16.0 |
| tvOS | 17.0 |
| macOS | 14.0 |
| visionOS | 1.0 |
| Swift | 6.0 |
FFmpeg arrives through MPVKit's dynamic xcframeworks — no FFmpeg source is redistributed here and nothing needs building. Dolby Vision RPU conversion goes through libdovi, which MPVKit already links.
A host that already ships MPVKit adds zero new binary dependencies. One gotcha
if that host vendors its own MPVKit fork: SPM derives a path dependency's identity
from its directory name, so the fork's directory has to be called MPVKit for it
to override the mpvkit identity PrismCore asks for.
FFmpegBuild reports the build that is actually loaded — a line per linked
library with the header version beside it whenever the two disagree, FFmpeg's own
version string, and the capability answers that change what this engine does with
a source: the eac3 encoder (without it the audio bridge cannot run), the AV1
decoder and whether this device has hardware for it, dialogue boost, and the GPU
deinterlacer:
print(FFmpegBuild.summary)
// FFmpeg n8.1.2
// libavutil 60.26.102
// libavcodec 62.28.102
// …
// eac3 encoder: NO — audio bridge disabled
// av1 decoder: libdav1d (hardware: no)
// dialogue boost: no
// gpu deinterlacer: none — CPU routeWorth a line in a host's launch log and worth attaching to a bug report: half of
what this engine decides is a question about the build rather than the media, so
"audio track is missing" is often unreproducible until you know which FFmpeg was
asked. A host that vendors its own fork should also assert FFmpegBuild.isABIMatched
once — a major version apart from the headers PrismCore compiled against means
the public structs this engine walks on every packet may be laid out differently,
which surfaces as wrong pixels and wrong timestamps rather than as an error.
PrismCore logs and reports that state but does not refuse to run on it.
FFmpegBuild.configuration carries libavcodec's full configure line for the
report.
prismcore-cli is a macOS tool in this package (not a product a host needs to
link) that runs the engine on one source from the command line:
swift run prismcore-cli probe <url-or-path> # SourceInfo, structure, routing verdict + reason
# (--json: the diagnostic report, alone on stdout)
swift run prismcore-cli serve <url-or-path> # loopback playlist URL for Safari / QuickTime
swift run prismcore-cli bench <url-or-path> # the startup checkpoint line a host logs
swift run prismcore-cli segverify <url-or-path> # decode every served segment on its own
swift run prismcore-cli validate <url-or-path> # Apple's mediastreamvalidator, if installed
serve and segverify take --report FILE and write the session's diagnostic
report there on exit, a failed start included.
validate is opt-in. It needs Apple's HTTP Live Streaming Tools, and without
them it prints a notice and exits 0 (unless you pass --require-validator).
The CLI does not replace a device run. The tvOS display handshake, Dolby
Vision on a real panel and Atmos passthrough can only be checked on hardware.
--help lists the options, and AGENTS.md Measuring explains how to benchmark
honestly.
PrismCore follows Semantic Versioning. Every public
declaration in Sources/PrismCore/ is the stability contract; internal types are
not. Since 1.0.0 that contract is a promise: breaking changes bump the major,
features the minor, fixes the patch — so the ordinary pin is the right one:
.package(url: "https://github.com/Wenzlik/PrismCore.git", from: "3.0.0")Hosts that archive through Xcode Cloud (or any CI with automatic resolution
disabled) should pin an exact version and keep their committed Package.resolved in
step with it. A from: range lets local resolution ride ahead to a newer tag while
the pinned file stays put, and then a device build and a TestFlight build are quietly
running different engines.
Every release is listed in CHANGELOG.md.
PrismCore is LGPL-2.1-or-later with an Application Store Exception (see
LICENSE and NOTICE.md). In practice, for the two cases
that come up:
Shipping a closed-source app, including on the App Store. Add the package, use
it, ship it. The exception exists precisely for this: LGPL section 6 asks that your
users be able to relink your app against a modified PrismCore, which a signed .ipa
cannot allow, and the exception releases you from that requirement for store
distribution. Two things are asked of you in return:
- If you modify PrismCore itself, those modifications stay LGPL and have to be published. Using it unmodified — the normal case — carries no such obligation for your own code.
- Say so somewhere a user can find it (an About screen, an acknowledgements page): that your app includes PrismCore, under which licence, and where the source lives. Aether's About screen is a worked example.
FFmpeg is a separate obligation, and it is not new. The exception covers PrismCore's own code only. FFmpeg reaches your app as dynamically linked frameworks under plain LGPL-2.1-or-later, which is satisfiable for a closed-source store app — that is what every FFmpeg-based app on the App Store already does — provided you embed them dynamically, ship the licence texts, and point at the build's source. If you already ship MPVKit, you are already doing all three.
Why LGPL rather than something permissive: Sources/PrismCore/Remux/EAC3Syncframe.swift
was written by reading FFmpeg's ac3_parser.c, because the placement of the Atmos
signal in addbsi cannot be derived reliably from the specification alone — three
attempts from the spec produced a confidently wrong answer. That makes the file
realistically a derivative of LGPL code, and NOTICE.md says so out
loud rather than hoping nobody looks.
Shipping, and pre-1.0 for the reason the version number suggests: the API is still free to move, and two headline claims are proven by construction rather than by a lab.
The test suite is 253 tests across 44 suites, hermetic apart from an opt-in
real-media harness. It exercises the playlists, the loopback server, the subtitle
renditions and the audio bridge's decode → resample → FIFO → encode chain, plus
end-to-end remuxes of synthetic fixtures: a two-language master (AAC eng + AC3 ces)
is served over the loopback, re-probed through libavformat's hls demuxer with both
codecs and both languages intact, and reaches .readyToPlay in a real AVPlayer.
Bitmap-subtitle OCR is verified against a real Blu-ray remux, where the served
segments read back the film's own SDH cues.
Being precise about the two headline claims, because "implemented" and "proven" are
not the same thing. Atmos: EAC3+JOC stream-copies (proved on a fixture — the
codec survives the round-trip), the rendition declares CHANNELS="16/JOC" (the rule
is unit-tested), and the served dec3 declares complexity index 16 on Dolby's own
demo file. Whether the objects reach a receiver over HDMI is a device question.
Dolby Vision: the whole signaling path exists and every rule is pinned by a unit
test, but a synthetic fixture cannot carry an RPU, so Profile 7 conversion is
established only as far as "the library links and the code runs". Both want a device
and real media, and that is what the Aether integration is currently buying.
The software path is exercised headless as far as it can be — the VP9 fixture decodes
to CVPixelBuffers of the right shape on a monotonic, source-anchored timeline, and
the pipeline runs demux → decode → stamp → back-pressure → enqueue against renderer
stand-ins, including the case where a stalled video renderer must not starve audio.
Audio track switching is proven the same way: the multi-audio fixture's tracks differ
in channel count (AAC stereo vs AC-3 5.1), so the renderer-side buffers show which
track is playing, and the video assertions show the switch never touched the picture.
Text selection is tested against embedded full/forced subtitle tracks, including
paused switching, Off, expiry and a backward seek. What no headless test can
assert is that a frame reached a display or an audio switch was inaudible on a
device. Caption rendering belongs to the host; frame-accurate seek remains open.
let session = try PrismCoreSession(
url: sourceURL,
display: .current(),
audioDelaySeconds: 0.150, // present audio 150 ms later
coordinatedHTTP: true // optional; finite HTTP(S) Range files only
)
let playlist = try await session.start()
let ranges = session.residentRanges
let routes = session.audioTrackDeliveries
let preview = try await session.cachedThumbnail(at: sourceSeconds)
let clip = try await session.cachedAudio(from: sourceSeconds - 10, duration: 10)residentRanges and cachedThumbnail use source timestamps, including a
nonzero source origin where present. A host must map these onto its player
timeline. These are video-cache ranges, not AVPlayer's loaded ranges or a
promise that a seek has no latency. Cache misses return nil without asking
the source for data; previews show the segment's opening picture. Fragments
above 64 MiB also return nil. The image cache holds at most 16 MiB / 32 images.
cachedAudio(from:duration:renditionName:) is the same contract for sound:
the resident audio of the playing rendition (renditionName is the master's
NAME; nil means the DEFAULT one) decoded to 48 kHz mono Float32, at most
30 s, with the audio offset applied so a clip is what was heard at that time.
It never fetches, seeks, re-anchors or arms a lazy rendition, so it is safe
to call while the title plays; nil means some of the interval is not on
disk, and an unknown renditionName throws CachedAudioError.unknownRendition.
audioDelivery summarizes available base routes, while audioTrackDeliveries
lists each source track, including an unavailable one. The host owns actual
AVPlayer track selection. A .bridged route describes its constructed pipeline;
its bridge counters show whether output has actually been produced. Zero
output during priming is normal. A terminal, completely drained silent bridge
raises AudioBridgeFailure.producedNoAudio through the session error path.
The software pipeline exposes .decoded only after its decoder opens.
Only the DEFAULT audio rendition and stream-copied ones are produced from the
first packet. A bridged track that is not the DEFAULT, and every dialogue-boost
level, is declared in the master and started by AVPlayer's first fetch under
its directory (a track switch re-anchors production at the demanded segment).
A sequential session — no keyframe plan, so no demand seam — leaves those
renditions out of the master instead of encoding them for the whole film.
audioRenditionProductions lists which was which, and
AudioRenditionProduction.summary(_:) is its one-line form for a startup log.
Audio delay is fixed at construction and carried by every clone. Changing it mid-title means a replacement session and a host-managed handover at the current position — see Changing a setting mid-title below. Both positive and negative offsets are bounded to two seconds; this does not add an AVPlayer transport controller to PrismCore. Audio delay can be changed while the title plays — it is a lip-sync control — and is preserved by fallback factories. Both positive and negative offsets are bounded to two seconds, and a non-finite value becomes zero. The two paths take a new value up differently, and each reports which:
switch session.setAudioDelaySeconds(0.2) { // remux path
case .pendingReanchor: break // accepted; in force at the producer's re-anchor
case .inForce: break // only before start()
case .unsupported: break // no plan (live, or no usable index): unchanged
case .sessionStopped: break
}
session.audioDelaySeconds // what is being SERVED right now
session.pendingAudioDelaySeconds // a request not yet in force, else nil
pipeline.setAudioDelaySeconds(0.2) // software path: in force when it has runOn the remux path the engine is serving fMP4 segments written with the previous
offset, and the offset moves audio dts — which cannot step backwards inside a
fragment the muxer is already writing. So the call asks the producer to
re-anchor at the playhead; at that re-anchor the new offset goes in force and
every segment written with the old one is discarded, so a later seek cannot
serve audio at the offset the viewer corrected away from. The picture re-buffers
while production catches up, and AVPlayer plays whatever it had already buffered
at the old offset first — the host should say so in its UI. audioDelaySeconds
names only what is in force.
On the software path the offset is applied where a decoded buffer reaches the renderer: the call flushes the audio renderer and refills it from the source at the playhead, so the new value is in force as soon as the call has run. The cost is a gap of decode-to-playhead time, not a re-buffer. The clock, the picture and the subtitles are untouched on both paths; this does not add an AVPlayer transport controller to PrismCore.
coordinatedHTTP is also available on PrismCoreEngine.open,
SourceProbe.open/openDetached, SeekPreviewService and software load(url:).
Enable it consistently for the readers that should share a server's request
budget. The default remains native FFmpeg I/O. The optional reader uses bounded
1 MiB Range requests and requires valid 206/Content-Range replies; it is not
for live streams, nested HLS playlists or Range-less servers. A redirected
request gets its destination's budget, drops caller headers on a cross-origin
redirect, and refuses HTTPS-to-HTTP downgrade. Thus an authenticated redirect
requiring forwarded custom headers should use the default transport or a final
URL supplied by the host. CDN performance has not been benchmarked.
To run the picture-based seek check on a Mac with AVFoundation rendering:
PRISMCORE_RENDERED_SEEK=1 DEVELOPER_DIR=/Applications/Xcode.app/Contents/Developer \
swift test --filter pictureMatchesPlayerClockAfterSeeksThe committed seek_clock.mkv has a reproducible generator beside it and runs
through a throttled, Range-capable test origin. The check reads actual decoded
frame numbers after seeks; it does not infer success from muxer timestamps.
A session is single-use: the served shape is decided before the first packet and
the output layout follows from it. "Play this differently" — the viewer turns
dialogue boost on, a smaller disk budget, a refused master — therefore means a
new session over the same source. makeSession(changing:) mints one from what
this session was built with, so the host does not have to restate it:
let boosted = try await session.makeSession { $0.dialogueBoost = [.medium] }
let playlist = try await boosted.start()
player.replaceCurrentItem(with: AVPlayerItem(url: playlist))
await player.seek(to: resumeTime)
await session.stop() // the caller's jobPrismCoreSession.Options carries everything the initializers take. Registered
external subtitles and the timed-text cue handler are replayed onto the
successor, and every option the closure leaves alone — the audio delay included
— is carried verbatim. sourceURL and httpHeaders are read-only: the replay
is what makes them part of a session's identity, and attaching one title's
captions to another file is not a clone.
Three rules, enforced rather than merely documented:
- Nothing is seamless. The successor starts from zero. There is no shared
playhead and no continuity of playback; the host replaces its
AVPlayerItemand seeks to where it wants to resume. - The caller still stops the predecessor. Nothing is stopped for you — at
the moment of the call the player may still be drawing frames off it. Stop it
as soon as the successor's playlist is loaded: every live session carries a
producer reading the source, a server on its own port, and its own
segmentCacheBytesbudget. - A session mints one successor; a second call throws
SessionError.alreadySuperseded. Clone the session you are playing, not the one you left behind. The successor never inherits the predecessor's work directory, which is what keeps two producers from writing the same segment names — and keepsstop()from deleting a directory somebody is still serving from.
makeMuxedFallbackSession() and makeMasterRejectionFallbackSession() are the
same mechanism with the option preset, and count as that session's one
successor.
Questions, integration help and bug reports: GitHub Issues.
Václav Zmrhal — zmrhal.cz
Written for Aether, where PrismCore is the engine that lets an MKV keep its Dolby Atmos and Dolby Vision instead of losing both on the way to the screen.
Built in close pair-programming with Claude (Anthropic); the commit log carries the receipts.
Copyright © 2026 Václav Zmrhal.
LGPL-2.1-or-later with an Application Store Exception. See
Integrating it for what that means in practice, and
NOTICE.md for the third-party components and their terms.
