Repository navigation
Per-run dollar budgets for Haystack pipelines and agents (no custom component) #13048
domondi1
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
One Haystack pipeline often serves many runs, and an agent run can make several model calls. This shows a way to give each run its own dollar budget with the built-in
OpenAIChatGenerator, enforced outside the pipeline.Point the generator at a local OpenAI-compatible gateway (Inferrail) and pass the run's id and budget per
pipeline.run:The first call of a run creates its budget. Calls that no longer fit get HTTP 402 before the provider (the generator raises
openai.APIStatusErrorwithstatus_code == 402; insidepipeline.runHaystack wraps it inPipelineRuntimeError), andinferrail work support-ticket-4812prints the run's cost afterwards. Tested with Haystack 3.2.0 and Inferrail 0.4.8 (one pipeline, two runs; the second run's call was refused at its budget).Limits: each call reserves an estimate (prompt plus
max_tokens), so setmax_tokens; only calls through the gateway are counted. Walkthrough: recipe. Disclosure: I maintain Inferrail (Apache-2.0).All reactions