Resource-aware Step Distribution Across Workers #4275
Closed
anthonyjoeseph
started this conversation in
Ideas
Replies: 1 comment
|
I've found a solution! It involves inngest 'connect', which makes sense since serverless setups won't generally have this issue I hadn't realized that each inngest 'connection' is able to set its own
This is perfect, and it's something I was able to base a solution on. Here's a repo with a full working example + video: https://github.com/anthonyjoeseph/inngest-worker-concurrency-cap-test |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
I originally asked this here as a feature request, but the response I got didn't solve my problem. I realized - maybe a discussion is more appropriate!
Is your feature request related to a problem? Please describe.
I have a step that requires a lot of memory, around 100mb, and worker machines that don't have much, 512mb. I'm worried that running this step more than twice in parallel will crash my workers
The concurrency controls are NOT a solution because they are scoped to functions, accounts or environments rather than individual steps, and are worker-machine agnostic. They're meant to avoid overwhelming external resources, but I'm concerned with internal resources
For example, I'm able to limit the maximum number of concurrently executing steps within a function w/ "Basic Concurrency," but this doesn't help in my case - I'd like to be able to execute 2 steps at once per machine, at max. Let's say I have two machines. If I set my limit to "4", there's nothing stopping the scheduler from executing all four steps on the same machine.
Deploying this step in its own function within a separate app (as suggested in the original thread) is also NOT a solution. Apps are helpful for isolating particular steps to a particular pool of machines, but they don't help control the usage of an individual machine. For example, let's say I have two machines assigned to this new app. If I assign 'concurrency: 4' to this new function, there's still nothing stopping all four executions from happening on the same machine in parallel
Describe the solution you'd like
My best idea at the moment is - I'd like to be able to specify, for a given step, a maximum concurrency 'per-worker' - the maximum number of times this step can be run in parallel on the same worker machine. In case this number is specified, I'd like that worker to be 'locked' - no other steps allowed to run in the background while the 'locked' step is running
I know this is slightly against the inngest abstraction - typically, inngest handles worker orchestration as a black box - but this seems common enough to be worth considering and hopefully simple enough not be too difficult to implement
Describe alternatives you've considered
Temporal offers something similar - docs
This is automatic, though, and doesn't have the kind of manual controls I think would be ideal
(Incidentally, there is a request for manual controls here - temporalio/temporal#8356)
I could spin up a separate message queue and integrate it with my inngest function, but that would introduce a lot of complexity and breaks the inngest abstraction much more than my proposed feature would imo
I think my best solution, in the meantime, is to implement some kind of file-system-based locking semaphore mechanism and throw an error when my machine's limit has been exceeded
Additional context
All reactions