docs/mimo: measured on an idle machine - #1820
Merged
Merged
Conversation
The numbers in docs/mimo.md were a first measurement taken while two other jobs held 13 of the machine's 16 threads (0.89 tok/s). On the same machine with nothing else running: 2.34 tok/s with 32 experts cached per layer, 2.95 with 64, 3.37 with int8 dense weights. MIMO_IDOT measured no faster, and the environment table now says so.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The measurement in
docs/mimo.mdwas taken while two other jobs held about 13 of the machine's 16 threads (0.89 tok/s at 4 threads). This replaces it with runs on the same machine with nothing else running.MIMO_IDOT=1MIMO_DENSE_BITS=8)Ryzen 7 PRO 8700GE, 64 GB, NVMe RAID, MiMo-V2.6-Flash, the CLI from a cold expert cache, a 33-token chat prompt and 128 greedy tokens; the command is in the doc. Per-token timestamps were taken from the CLI's stdout, which is flushed per token, so the doc also gives the rate over the last 64 tokens.
Each run's load and busiest process were recorded. A run that overlapped other work was discarded and repeated; the two 16-thread runs agreed within 1%.
MIMO_IDOTmeasured no faster here, so its row in the environment table no longer says "faster". Docs only.