chore(container): update image ghcr.io/ggml-org/llama.cpp ( e755b93 → 0c09831 ) - #2289
Conversation
e6c6b3f to
12a4abc
Compare
8bbc41a to
bff14f1
Compare
bff14f1 to
f2185d1
Compare
6d046c5 to
409950b
Compare
409950b to
6956c40
Compare
1Solon
left a comment
There was a problem hiding this comment.
Reviewed exact digest-only diff at 6956c40 and upstream 67a17c17caa95742186f8b1ecadd1b5abd6d5ebb...9dcf84e5ae2718947188b539aab8b9c2b15d3ba1 (78 commits). Qwen QKV/MTP changes preserve separate projections; GDN normalization is corrected; recurrent-layer metadata is converter-side; no required current-model/cache/CLI migration found. Router LRU changes do not apply to this fixed-model invocation. Inherits verified 34Gi request/limit from #2293. Server dry run passed. Proceeding with sequential rollout and loopback inference/OOM checks.
|
Post-merge verification passed at 0f4291b: live imageID sha256:0c09831866497d97802a57a36c4010253d3d43e692e521b947eec5fd4ef0af4d, fingerprint b10853-9dcf84e5a. Pod, Model, InferenceService and Flux Kustomizations ready. Request/limit remain 34Gi; memory.max=36507222016, observed peak=19827515392 bytes. Two loopback chat tests returned OK, both 2/2 MTP drafts accepted; repeat reused 13 prompt tokens. Health OK, zero restarts, all memory.events counters zero. kube5 Ready without memory pressure; NVIDIA GPU reporting works. Startup included unexplained nonfatal ERROR: init 250 result=11 and duplicate n-gpu-layers deprecation warning; neither prevented loading/inference. These are smoke checks, not a full-cache or long-context stress test. No storage/database changes or rollback performed. |
This PR contains the following updates:
e755b93→0c09831Configuration
📅 Schedule: (UTC)
🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.
♻ Rebasing: Whenever PR is behind base branch, or you tick the rebase/retry checkbox.
🔕 Ignore: Close this PR and you won't be reminded about these updates again.
This PR was generated by Mend Renovate. View the repository job log.