Repository navigation
RFC: a fixpoint that settles #617
Description
Activity
- changed the title
[-]RFC: a fixpoint that settles — recursive types by origin, monotone rules from ⊥, shared types, ordered rounds[/-][+]RFC: a fixpoint that settles[/+]on Oct 8, 2026 A few specific questions for the people whose work this builds on:
- @thomasklemm: S2 and S3 build on A type the fixpoint carries between rounds stays bounded (#518) #584 and Fixpoint bound follow-ups: record the Rust gap, pin that the fixpoint settles #589, and on Stabilize harvested returns that oscillate on Untyped/Var #521 and Narrow production typing rounds with a call-graph dirty set #524. Does finding 3 under Details, the non-monotone rules, match the oscillation you traced on Mastodon and Discourse? And should S3 share a design with RFC/spike: evaluate change-driven fixpoint scheduling within one fresh whole-program analysis matz/spinel#7237 (question 4)?
- @dai199: S2 changes
harvest_return.rs, where A recursive method's harvested return stops unrolling; a gap cause is the same on every run #528's cut lives. Where recursion is certified, it would store a recursive reference instead of cutting tountyped. Does that fit your plans for that file? - @eddygarcas: S2 needs every pending join to be commutative, as Make a module parameter's call-site join order-independent (#209) #565 made the module parameter join. Does extending that to all pending values look right to you?
- @tobi: S1 extends Make
roundhouse checkfast on large apps #171's sharing from constant scopes to the types themselves. On the large app from Analysis grows without bound on a helper that returns a recursive array (since #477) #518, the prototype completes the check at 3.06 GiB with no size bound, which may help on codebases like the one behind Meta PR 2: fixes from running roundhouse over a large internal Rails codebase #503. - @rubys: questions 1, 2 and 5: how to stage S1's migration, whether to split
untypedinto pending, gradual and unresolved, and where the proof and corpus should live. - @matz: question 4. S3 and RFC/spike: evaluate change-driven fixpoint scheduling within one fresh whole-program analysis matz/spinel#7237 face the same scheduling problem from opposite sides: a monotone analysis here, and Spinel's deliberately non-monotone pipeline there. S2's case for accumulation starts from the failure matz/spinel@9332663616 records.
A staged diff on current main, one commit per stage, follows in this thread when it's ready.
Thanks for this. It's a careful piece of work, and the diagnosis matches what I saw on Mastodon: when ccbafad made rounds cheap, the flipping was still there, and my attempts at a "no downgrade to
untyped" rule and a "no new facts" stopping rule didn't settle it either. Your explanation of why each half fails without the other is the clearest account of that I've read.On the questions you put to me:
S1 staging, and splitting
untyped: you've got this. You know the migration's shape better than I do. Splitting pending, gradual and unresolved is the lesson I took from 68f4d82, and I'm glad to see it made explicit. Sequence it however works best for you. The gates you've written (byte-identical output, a green suite, a paired measurement against #584) are the right ones.Where the proof, corpus and oracle live: your call too. Whatever a CI gate depends on probably wants to be in this repo eventually, but there's no rush.
A few things I'd be curious whether you've considered:
- Identical results run to run as the first gate. The three large-app runs giving different digests stood out to me more than the precision or timing numbers. Have you considered making "same answer every time, in any order" the bar S2 has to clear before anything else is weighed?
- Measuring each stage in roundhouse's terms. Settling and least solutions are the analysis's own yardsticks. Have you considered also reporting each stage by error counts,
checktime on Mastodon and Discourse, anduntypedin the generated code? Those are what tell us more of Rails transpiles. - Where recursive types would pay off. Generated code is byte-identical and typed recursive emission is out of scope, so the payoff for the transpiled output is deferred. Have you considered naming one concrete target, say typed recursion in the Rust or Crystal output where Fixpoint bound follow-ups: record the Rust gap, pin that the fixpoint settles #589 recorded gaps, as the reason the later stages are worth it?
- S3 earlier. On its own the ordered worklist took about half the time on the large app, and that's the piece that speaks most directly to Mastodon's cost. With the shadow mode you describe (checking that every skipped method really wouldn't have changed), have you considered whether it could come before S2 rather than after?
- The
untypedsplit as its own step. Have you considered landing the three kinds first as a tag that changes no behavior, with byte-identical output, and changing the join rules afterward behind a flag?
Looking forward to the staged diffs.
Thanks, that's generous, and so is the latitude on S1 and on where the proof, corpus and oracle live.
Both rules you tried fail for reasons the RFC predicts:
- "No downgrade to
untyped" freezes an answer once it looks informative, so when a value genuinely changes, the stale answer stays. matz found the same in Spinel: the rule "froze a return that had genuinely become wrong" (matz/spinel@9332663616). - "No new facts" stops when a round adds nothing. But a rule that isn't monotone can also stop producing a fact it produced before, and that fact stays stored, stale. So "no new facts" means settled only when every rule is monotone (learning more never withdraws a conclusion) and the loop starts from nothing known yet. The Details block "Why accumulation can work now, after two failed attempts" works this through.
Each of your five suggestions changes the plan, and two come with a caution:
-
Same answer every time, in any order, as the first gate. Adopted, with a correction to the RFC first.
- Run to run, the large app's digests differ after the loop, not inside it. Under two different hash seeds, the state the loop settles on is identical; only the final expansion differs. That step unfolds recursive references into ordinary types for the emitters, and it walks HashMaps in hash order. A sorted walk is the planned fix.
- In shuffled order, every difference traced so far on the public apps comes through one of two rules.
decide_harvested_return(from Stabilize harvested returns that oscillate on Untyped/Var #521) chooses between a method's stored return and its newly harvested one, and several of its branches keep whichever was written first, or last.unify_param_ty, when both values are pending, keeps the later one, so HashMap iteration order can decide the result. - The plan: each rule gets an order-independent replacement, a join that gives the same answer whichever value arrives first. Repeated and shuffled-order runs check it before precision or timing is weighed. Main's own output already varies by a few lines between runs, as thomasklemm noted in #518, so the same replacements should help main too.
-
Each stage in roundhouse's terms. Adopted, with one caution. Every staged diff will report, next to the analysis's own measures, errors by kind,
checktime on Mastodon and Discourse, anduntypedin the generated code, counted in the.rbssidecars Spinel reads. The caution: both counts can mislead when a rule drops anuntypedarm that a value really has. The sidecar printsString | untypedas plainuntyped, so dropping the arm improves the count while making the type wrong. Errors mislead the same way: in the RFC'sjoin_untypedreproduction, current main reportsno known method `bit_length` on Stringfor a call that succeeds at runtime, because the dropped arm left onlyString. So each report will pair the counts with the runtime oracle, which checks inferred types against values recorded while the program runs. -
A concrete payoff. Adopted. The target is the normalizer from Analysis grows without bound on a helper that returns a recursive array (since #477) #518, transpiled with a real recursive type, so the generated code compiles and prints what Ruby prints. In Rust, where Fixpoint bound follow-ups: record the Rust gap, pin that the fixpoint settles #589 recorded the gap, that means an enum, with
case value when Hash / when Arraylowered to amatchon its variants. In Crystal, a recursive alias expresses the type directly.There is also a nearer payoff, found today. Main's test
a_result_merged_back_into_its_own_parameter_settles_bounded(inrecursive_type_bound) settles only because main drops the|k, v|block flow insort.to_h. Restoring that flow on main, a fix that only adds information, makes the test fail: its loops run to the cap. The prototype, given the same fix, settles the program. On main, a sounder type and a settling loop exclude each other here; the new fixpoint gives both. -
S3 earlier. Yes, with S3 split in two:
- The first half keeps today's round order and skips re-typing a method when nothing it read has changed since its last typing. A typing depends only on what it reads, so the same inputs give the same output, and skipping should change no answer, even under today's rules. Shadow mode re-types every skipped method anyway and checks that claim. This half can follow S1.
- The second half visits methods in dependency order. Under today's order-dependent rules that changes answers, so it waits for the replacements in (1).
One caution: the "about half the time" came from the full worklist, which also reordered, and so changed answers. How much of that gain the first half keeps isn't measured yet.
-
The
untypedsplit as its own step. Adopted. The three kinds land first as a tag that changes no output, so byte-identical output is the whole review. The join rules and recursive references follow, each behind its own flag, switched on together, since each half fails without the other.
The staged diffs will follow in this thread, each measured as in (2).
- "No downgrade to
@bunnykong Thanks for laying out how this touches #528.
Short answer: yes, it fits. #528's cut was a stopgap for one symptom, chatwoot's
sanitize_mailbox_valuedoubling every round. It decides "recursive" from the slot's own history, and a recorded origin is a better witness. I have nothing in flight onharvest_return.rsthat this would collide with.I built
phase-c2.diffonb28b17b6next to current main and tried a few small shapes. Below are three suggestions, plus one finding that the RFC doesn't cause.1. One merge decision per slot.
decide_harvested_returnis currently the one place that says what a write to a harvested return does. In the prototype,fold::join_retsjoins on top of it afterwards. So a reference-mode slot goes through #521's history-dependent stabilize first and the monotone join second. I'd rather S2 replace those rules insidedecide_harvested_returnthan add a pass after it.2. The Campfire case in the module header. It records that a full lattice join was tried and rejected, because
Untypedfollowed byNilcollapsed Campfire's URI helpers to bareNil. That is the pending/gradual confusion again. The split should fix it, but only if those helpers'Untypedis classified as gradual or unresolved rather than pending. It would make a good regression test for the producer-side split.3. Gates for this file. The
insert_recursive_return_*tests pin the cut. Under the fold they need SCC-keyed equivalents (thesanitize, flattened-union and record shapes). Chatwoot'scheck --continuetime and memory should be a gate too, since that was the hang #528 fixed (over 20 minutes and 30 GB before). For scheduling, a work-count check in the style of Spinel'smake scale-testcould help: generate the chain at two sizes and compare typings, not seconds. Your chain-64 runs in about 0.04 s here. On the prototype it settles with 128 typings (2 per method); main leaves 39 of the 64 returnsuntyped. A count like that doesn't depend on the machine and is cheap enough for the defaultcargo testjob, next totests/recursive_type_bound.rs, which already guards time and memory for the cycle shapes.On the cut outside certified components: on current main I couldn't make it misfire on a plain chain. I tried six methods each wrapping the next in an
Array, defined in either order, on the class side and the instance side, and every one was typed exactly. So I don't mind when the cut goes; reportingharvest_untie_cutper stage would show it becoming unused.A class/instance finding the RFC doesn't cause. I checked whether
def self.call(*a) = new(*a).callwould read as a self-call and putcallinto reference mode, since the call-graph edges are keyed by(class, method). It doesn't: the prototype reportsrec_methods: 0for that shape. But the same shape already failed earlier, on bothb28b17b6and main. Dispatch checkedclass_methodsbeforeinstance_methodson the sharedTy::Class, soGreeter.new(x).callresolved to the class-sidecall, whose body then resolved to itself and stayeduntyped. Forem writes this service-object shape in 107 files. #630 fixes it: across the seven public apps, warnings drop by 614 and chatwoot loses 19 errors; the 30 errors it adds are existing gaps that a return type now reaches, detailed there. The side ambiguity lives inTyitself, not only in the params and slot keys that #551's follow-up and S4 mention, so it's worth knowing before S2's join meets those slots.Thanks for building
phase-c2.diffand trying shapes against it, and for confirming that S2's change fits your plans forharvest_return.rs. All three suggestions are adopted.1. One merge decision per slot. Agreed. S2's order-independent join will replace the rules inside
decide_harvested_return, and the separatefold::join_retspass goes away. The function is central to rubys's first gate, the same answer in any order: every shuffled-order difference traced to a method on the public apps passes through it, because several of its branches depend on the order of writes. But the join alone won't meet the gate. In a first trial of the order-independent joins, Campfire gives the same answer in any order, while Chatwoot, Mastodon and Discourse still don't: some rules for typing method bodies aren't monotone yet, and which slots become references can depend on the order. The staged diffs will report each source.2. The Campfire case. Agreed on the diagnosis, and it becomes a regression test for the
untypedsplit. Your caveat held: a trial that treated every unknown return as not computed yet, failed method lookups included, brought the collapse toNilback. The test fails under that trial, and passes on main and when only genuinely pending returns are treated that way. Telling those apart needs the provenance tag, which is why the tag lands first.3. Gates for this file. All four go into the plan:
- SCC-keyed versions of the
insert_recursive_return_*tests, for thesanitize, flattened-union and record shapes; - Chatwoot's
check --continuetime and memory at every stage, since that was the hang A recursive method's harvested return stops unrolling; a gap cause is the same on every run #528 fixed, alongside Mastodon and Discourse; - a work-count test in the default
cargo testjob, next totests/recursive_type_bound.rs: the chain at two sizes, compared by typings rather than seconds, with the RFC's chain-64 as the first case; harvest_untie_cut, reported per stage.
On the cut outside certified components. Your six-method chains show that on main the cut leaves short chains alone. The RFC's finding 2 is a limit in principle: a rule that sees only a slot's own history can't tell a recursion from a long enough chain whose history grows the same way, one level per round. A recorded origin can. The per-stage counter will show whether the cut stops firing, and so far it hasn't: on the prototype it still fires on four of the five public apps, and a first trial of the order-independent joins makes it fire more often. The staged diffs will explain why.
On the class/instance finding. Thanks for tracing it and for #630. The side lives in
Tyitself, not only in the parameter rows that the RFC's S4 note mentions. Once #630 lands, thedef self.call(*a) = new(*a).callshape becomes one of S2's tests, checking both sides' returns and that neithercallenters reference mode.All of this goes into the staged diffs.
Reacted by daiki tagami- SCC-keyed versions of the
- added a commit that references this issue
on Oct 8, 2026 The staged diffs are up: six commits on main
194f26cf, one per stage, with theuntypedtag as its own step, so S2 is three commits. With every flag off, each commit gives main's output: emission is byte-identical on all 105 fixture × target pairs, the library tests pass at every commit, and the default suite passes at the tip. The PRs will be rebased onto current main.Each stage in roundhouse's terms on Mastodon · Discourse · Chatwoot, against main at
194f26cf. Errors were identical across three runs; times are medians of three interleaved runs.Stage Errors checksLoops main 1,583 · 3,136 · 958 9.1 · 20.9 · 4.3 production at the cap on all three; absorb too, except on Chatwoot S0. Canaries, opt-in as main as main as main S1. Shared types as main 8.5 · 19.7 · 4.0; peaks 13–16% lower as main S2a. The untypedtagas main 9.1 · 20.2 · 3.9 as main S2b. Joins, alone +11 · −4 · +1 8.7 · 21.5 · 5.8 every loop at the cap S2c. Recursive references +1 · −5 · as main 9.2 · 18.9 · 5.0 Mastodon and Discourse settle; Chatwoot's production doesn't S3. Dependency-ordered worklist as main · −13 · as main 5.5 · 16.9 · 3.0 every loop settles, on all five public apps S2b is measured alone only because the split asks for it; it ships switched on with S2c. Its 11 extra errors on Mastodon are 7 dispatch and 4 binop. S3's 13 fewer on Discourse are 10 dispatch and 3 binop.
untypedin the generated code, the share of signature positions in the Spinel sidecars, pooled over the five public apps: 40.89% of signature positions holduntypedon main, S1 and S2a; 41.19% at S2b alone; 40.99% at S2c and S3.Precision, fully typed with the runtime oracle beside it: 69.02% of expressions are fully typed on main and 68.83% at S3, pooled over the five apps, and the share holding
untypedgoes from 14.34% to 14.55%. Beside that, of 17,820 values recorded at runtime across 53 traces, S3's types reject 6,699 and main's 9,224. Most of the difference is in the recursive reproductions; the settling reproduction below still needs the flow fix (8 rejected at S3, 6 on main).On S3 earlier: the S3 commit is the whole worklist, so its times include the reordering. The first half, skipping unchanged methods in today's order with a shadow check, isn't on the branch yet, and how much of the gain it keeps isn't measured.
The large app from #518 (aggregates only), at S3: production settles in 6 rounds and absorb in 3, and two runs agree in every digest. Errors are 3,598 against main's 3,675, 1.47 points more expressions hold
untyped(24.66% against 23.19%), and peak memory is 3.0 GiB against main's 4.0, mostly from S1. It is still slower: 47 and 66 s on a loaded host, against 39 s for S2a, which gives main's answers with S1's sharing, in the same series.The same answer every time:
- Run to run, it holds. Two runs at S3 agree in every digest on all five public apps and on the large app; S2c now walks the final expansion in sorted order.
- In shuffled order, the branch can't check it yet; S0's shuffled-order control comes next. On the RFC's prototype, a trial with the one merge described below removes the order dependence in merges and shows what remains. None of it is a merge. What remains includes body rules that settle differently depending on order (on Chatwoot,
chunkends asArray[Int]in one order andArray[untyped]in another), and the choice of which slots become references.
The two payoffs from the earlier reply, now public reproductions:
- On main, a sounder type costs settling; S2 and S3 with the fix give both. Main settles but rejects 6 of 12 values recorded at runtime, because it drops the
|k, v|flow insort.to_h. Restoring the flow on main runs production and absorb to the cap. S2 and S3 alone settle but reject 8 of 12. With the corrected flow rules on top (branchfixpoint-sound, behindRH_FOLD), every loop settles within 3 rounds and none of the 12 are rejected. - Typed recursion in generated code. On the RFC's prototype, behind a flag, a recursive reference prints as a Rust
enum(withcase … when Hash / when Arraylowered to amatchon its variants) and as a Crystalalias. For three of the four shapes Fixpoint bound follow-ups: record the Rust gap, pin that the fixpoint settles #589 recorded,cargo checkgoes from 3–5 errors to none,crystal buildpasses, and each program renders what CRuby renders, byte for byte. The class-method cycle isn't covered yet, and the emitter handles only these walk idioms.
dai199's suggestions:
- One merge decision per slot isn't on the branch yet: S2b still joins after
decide_harvested_return. A trial that makes it the one merge settles and repeats run to run. But it adds dispatch errors on Chatwoot, Forem, Mastodon, Discourse and the large app, on calls thatT | untypedused to absorb, so it waits until the error report counts those apart. - The Campfire case has been a test since S2b, and it caught a real bug. S2b's join dropped a gradual
untypedfrom anil-only core, which turned the URI helper into a method that always returnsnil. S2b now keeps thatuntyped, and the test passes at every stage.
dai199's gates for
harvest_return.rs- The shapes the
insert_recursive_return_*tests pin settle by reference under S2c, and the cut never fires on them. tests/chain_typing_count.rs, in the defaultcargo testjob, settles chains of 32 and 64 methods with two typings per method, every link typed. Main leaves 7 of 32 and 39 of 64 returnsuntyped.harvest_untie_cut, Mastodon · Discourse · Chatwoot, falls from 44 · 195 · 83 on main to 18 · 58 · 46 at S3. It no longer fires on the recursive walkers it was written for (Chatwoot'ssanitize_mailbox_value, Discourse'smap_jsonand others). The cuts that remain aren't recursion: most withdraw a known return when anuntypedarm joins it, and the rest cut a pending placeholder. The provenance tag tells both apart from recursion, so the next step is to cut only strict nesting of an informative return.- Chatwoot completes in under 8 s and 0.6 GiB at every stage.
Still open:
- the large app's speed under S2 and S3;
- the same answer in shuffled order;
- the
name_collisionreproduction, where a class method and an instance method of one name share a parameter row. It needs rows keyed by side, the follow-up named in A call through x.class does not feed the instance method of that name #551.
Next: S0 and S1 as the first PRs, with S1 split into shared payloads and the single walk over each shared node, each rebased and measured on current main. Corrections to the original proposal are listed at the top of the issue.
Since the last update, the first two PRs are open, a prototype addresses the large app's speed under S2 and S3, and errors are now matched call site by call site, which corrects how the last update read Discourse's 13 fewer errors. Two results are setbacks: at least one precision gain is unsound, and freezing part of the structure changed answers.
The PRs. #657 adds S0's opt-in canaries. #656 is the first half of S1: a copied type shares its pieces instead of duplicating them. Both pass CI.
The large app's speed, left open last time. A profile of S3 on the large app finds joins in 17% of the analyzer's samples, resolving recursive references in 15% and comparing types in 8% (the shares overlap). A prototype behind
RH_ARENAgives each type two identities:- an exact one, which keeps record field order and the
untypedtag; - a semantic one, which agrees with
Ty::eq.
The prototype remembers each union by its ordered pair of exact identities, so a repeated join is a lookup, and it resolves references by semantic identity. No typing rule changes.
With every S2 and S3 flag on, the large app's
checktakes 34.0 s with the prototype, against 43.0 s without it and 41.7 s on main194f26cf. Last time's comparison was with S2a; this one includes main. These are medians of three interleaved runs on a loaded host, and in every run the prototype beat both. On Mastodon and Discourse it takes another 6% and 8% off S3's time (Chatwoot wasn't in this series). Peak memory on the large app is 3.1 GiB, 2% above S3 alone and still well below main's 4.0, mostly thanks to S1.All ten S0 digests, the errors by kind and the diagnostic content are identical with and without the prototype on the five public apps and the large app, and the default suite passes. With the S2 and S3 flags off, the union memo alone still takes 3–5% off
checktime on Mastodon, Discourse and the large app, so it can follow S1 as a PR of its own; resolving references by semantic identity goes with S2c.A bug in the staged branch's sharing, found by the prototype and fixed
S1's interner reuses any live value that is equal apart from the
untypedtag, so a value could come back carrying another value's tag. The join memo, which lasts for a single join, compares the same way. No typing rule reads the tag yet, and with the interner's fix every S0 digest, error count and loop outcome is unchanged on the five public apps and the large app (the digests, likeTy::eq, ignore the tag). The tags themselves change on all six: on Mastodon, theuntypedleaves tagged gradual at the end of the run fall from 12,902 to 5,370, and the difference is now tagged unresolved. The join memo's fix passes its own tests; its run on the apps comes next. The narrower cut proposed last time will read the tag, so both fixes ship with S2a, where the tag first appears. #656 has neither the interner nor the memo, so it isn't affected.Errors by kind, classified. A census now matches each failing call site between main and S3:
- S3 adds no error on any of the five public apps.
- Of Discourse's 13 fewer errors, 3 are binary operations whose operands gained an arm that makes them compatible; whether that arm occurs at runtime isn't checked yet.
- The other 10 are dispatch errors whose receiver now includes
untyped, which hides them rather than fixes them.
So the top section's "as main or fewer" holds only for the count; the section now says so, and the gate counts hidden errors on their own line from now on. The census needs name-level output, so it runs on the public apps only; the large app's dispatch +10 and unsupported −85 remain counts by kind.
Precision, expression by expression. Last time S3 was 0.19 points below main in fully typed expressions. Matching every expression between the two, about a million on the five public apps, shows S3 losing precision on 4,823 and gaining it on 2,334. About half of the losses already hold
untypedbefore recursive references are expanded, a quarter pick it up during expansion, and the rest hold an unresolved type variable. Switching rules off one at a time puts roughly a third each on the worklist's order, on S2c's reference slots, and on S2b's joins and all-arms binding, counting a loss that several switches recover only once.Some causes are specific and have small fixes. Two are prototyped behind flags but not yet measured on the apps:
- a class test that succeeds but still lets
untypedinto the branch; Array(…)applied to a recursive reference, which treats the reference as a single value: in a focused test, an array of strings comes out asArray[Array[String]], a wrong type rather than a vague one.
A third, the class-versus-instance dispatch, is already fixed on main by #630, which the staged branch predates.
At least one gain is unsound. In Discourse's
app/jobs/base.rb, S3 types@dataas a hash ofIntegervalues after both aStringand anIntegerare stored in it, because the join over its[]=writes loses the earlier contents. Only a sample of gains was checked, and no runtime trace covers the apps, so the oracle couldn't catch this. The fix is to keep the prior contents in that join, and traces of app code would let the oracle catch the next one.The same answer in shuffled order. It can now be checked. On a follow-up branch, S0's canaries gain a seeded shuffle of the worklist and a digest of the structure the fixpoint works over, meaning which slots exist, who writes each one, and which become recursive references. At S3 today, that structure differs between schedules on four of the five public apps and changes during the run on all five. So there is no single fixpoint for the order to reach, and repairing the rules alone can't give the same answer in any order. This is also where the analyzer departs from the lab's Lean model, whose slots and constraints come from the program before solving.
A first attempt at fixing the structure changed answers. It builds part of the structure before typing: which methods are recursive, and the call graph between them. On the four public apps measured so far, that part now stays the same from start to end and across schedules. The rest is still created while typing, and it still differs between schedules on three of those four apps. Freezing only part of it also changed answers: errors appear on all four apps (on Forem, dispatch errors rise from 177 at S3 to 219), and some are hidden on Chatwoot and Mastodon. So this step stays local until the rest is declared too.
Still open: the same answer in shuffled order, and the
name_collisionreproduction, which waits for parameter rows keyed by side, the follow-up named in #551.Next:
- the second half of S1 once A cloned type shares its children instead of copying them #656 lands, then the union memo as its own PR;
- S2a as a PR, with both fixes for the tag;
- the unsound
[]=join fixed, and the precision fixes measured on the apps; - the staged branch rebased onto a main that includes A call on an instance answers the instance method of its name #630;
- dai199's one merge per slot, run through the census;
- the rest of the structure declared before typing.
- an exact one, which keeps record field order and the
The research behind this RFC is now public in #669, a draft not proposed for merge in this form: six open problems, a check for each, the established facts with their commits, and every attempt so far, dead ends included. The problems are hard and still open, and progress will be faster with more people working on them. Anyone is welcome to take one on: the README shows how to pick a problem, run its check on the pinned public apps, and report back here.
@dai199, your suggestion of a single merge decision per slot has now been through the census, as the last update planned. The trial is on
fixpoint-onemerge, behindRH_DET, off by default. It bundles the one merge indecide_harvested_returnwith normalized joins, so its effects aren't the merge's alone:- It settles, gives the same answer every run, and raises the pooled share of fully typed expressions from S3's 68.83% to 70.77% on the five public apps, above main's 69.02%.
- Against S3, it also adds 107 errors and removes 20. A review against the source judged 80 of the 84 new dispatch errors impossible and four unclear; none is confirmed real.
- In 59 of them, the receiver is
nilalone where the source builds an object. At S3, 53 of those receivers held an unresolved type variable and 6 anuntypedarm. The cause at the producer isn't traced yet, but because those receivers now count as fully typed, part of the gain is unsound. - In 9 more, Mastodon's
can?surfaced once its receiver lost anuntypedarm.Userdelegatescan?to its role, and the analyzer doesn't model that delegation;untypedused to hide the gap. - The other 16 dispatch errors have mixed causes, and the 23 errors of other kinds weren't reviewed.
One question, where a link to the lines would be enough: at S3, most of those receivers held a pending value beside
nil, and in the trial onlynilis left. Is there a place indecide_harvested_return, or in the join, where a pending arm next tonilcan be dropped, and should it be kept as unresolved instead? It resembles the Campfire case you raised, with a pending value where Campfire had a gradual one.Your recent PRs already appear in the research. #630 removed one of the causes of lost precision, and #634 changed main's
to_hpair reading, which the lab's settling reproduction depends on; re-checked on current main, that reproduction still holds.@bunnykong, on your question. Short answer: yes, but commutativity is not enough on its own. Every join that feeds a carried slot needs to be a real join-semilattice: commutative, associative and idempotent, with one operator per slot and with pending as its identity. On main, several joins fall short of that in ways #565 didn't touch. And as your structure digest already shows, the join laws give a schedule-independent answer only when the transfer rules are monotone and the slot set is fixed before solving. So I'd treat the joins as necessary, not sufficient.
What #565 did, and how far it carries. It made
unify_param_tysymmetric by orderingVar < Untyped < concrete, checkingVaron both sides beforeUntyped. It also sorted the includers by class id before folding, but that sort only gives the same answer run to run. It does nothing for a shuffled worklist, so it can't stand in for a commutative join. The ordering part does generalize, and it's cheap, becauseunion_ofalready canonicalizes. One caveat about our own change: absorbingUntypedbelow a concrete type is right for pending and unresolved values, and wrong for gradual ones. It drops a real arm, the same failure as yourjoin_untypedreproduction. Once S2a's tag exists, the rule should read: pending and unresolved are identities, and gradual stays as an arm. Campfire'sstart_new_session_for(user)would still getUser, because thatuntypedis a registry gap (unresolved), not a dynamic value.Joins on main that aren't semilattice joins (at d4f4076; each one confirmed with a scratch unit test):
unify_param_tywith two pending values keeps the later one:unify(Var(1), Var(2)) = Var(2), butunify(Var(2), Var(1)) = Var(1)(mod.rs#L6876-L6881). This is the case you traced.unify_param_tyisn't associative acrossunion_of. It absorbs a bareUntypedbut not anUntypedarm inside a union, sounify(Int, Str | untyped) = Int | Str | untypedwhileunify(unify(Int, Str), untyped) = Int | Str. The answer depends on whether a producer joined first. The fix is to define the join on the flattened arm set and classify each arm (mod.rs#L6865-L6908).union_ofandunify_param_tydisagree aboutVar.union_ofkeeps it as an ordinary arm, whileunify_param_tyandwiden_hash_ivar_valuetreat it as bottom. Sowiden_hash_ivar_valuegivesHash[Str, Int | Str]when the{}seed comes first, andHash[Str | Var, Int | Str | Var]when a[]=write comes before it (mod.rs#L7796-L7821, body/mod.rs#L2665).widen_hash_ivar_valueignores a[]=write when the ivar isHash | Nil, or anything else that isn't a bareHash(mod.rs#L7806). A nullable hash keepsHash[Str, Str] | Nilafter anIntwrite. That is not an ordering bug but a lost write, close to the unsound@datacase on your soundness frontier.- Record order breaks canonical unions.
Row.fieldsis anIndexMap, so==ignores field order, butcmp_rowcompares in insertion order (ty.rs#L550-L563).union_of({a: Int, b: Int}, {a: Str, b: Int})andunion_of({a: Str, b: Int}, {b: Int, a: Int})come out as unions with different variant order, and the two aren't==.push_union_variantsalso keeps whichever field order arrived first (body/mod.rs#L2762). Records also join as separate arms rather than field by field, unlike Hash and Array. That keeps the laws, but it is worth deciding on before S2c's record shapes. - The size bound runs after every pairwise join (
bound(unify(slot, obs))at mod.rs#L5203 and #L5417, the ivar merges at L7611/L7624/L7643, and harvest_return.rs#L156).bounddoesn't distribute over the join. With a 10-deepArrayand two 300-classHashvalues, one grouping ends asuntypedand the other keeps the deep array. The fix is to bound once per slot per round, after the whole fold, or to treat the bound as a widening applied at one fixed point in the schedule. With S2c's references it should almost never fire, but a backstop that fires in an order-dependent way breaks the gate whenever it does. decide_harvested_returnis replacement, not a join (harvest_return.rs#L115-L142). It keeps the last write when the cores differ. For example,Str | untypedthenIntgivesInt, and the reverse order givesStr | untyped. It also has an untie step that depends on history. You and @dai199 already have this one infixpoint-onemerge. The only thing I'd add is that switching it from replacement to accumulation is safe only once pending and unresolved are identities, or stale arms from early, less-informed rounds stick for good. On main,strip_unknownmapsVartoUntyped(ty.rs#L447-L460), sostabilize_untyped_return_oscillationturns a pending arm besidenilinto a gradualNil | untyped(harvest_return.rs#L42-L55). That may bear on your question to @dai199 about pending arms next tonil.
union_ofitself is fine: its lattice-law tests (body/mod.rs#L3873-L3968) hold, apart from the record field order above.Cost. Sorting fold inputs costs O(n log n) per fold and is negligible, but it only buys determinism from run to run. Normalizing on the arm set costs about what
union_of's canonicalization already costs, and your union memo covers the repeats. Comparing record rows with keys in sorted order adds a sort per comparison, unless rows are canonicalized when they're built. If emitted field order matters for a target, the canonical form should keep a stable source order rather than arrival order.Offer. I can take items 1 to 6 as one PR against main, outside
decide_harvested_return, which stays withfixpoint-onemerge. The PR would contain:- a single join for parameter slots, defined over the classified arm set;
- consistent
Varhandling inunion_ofandwiden_hash_ivar_value; - records compared without regard to field order;
- the bound applied once after each slot's fold;
- generated lattice-law tests (commutative, associative, idempotent, pending as identity) over every carried-slot join, next to
union_of's tests andtests/concern_param_determinism.rs.
Before S2a's tag lands, it would use today's
VarandUntypedand leave gradual absorption alone. After S2a, it would switch to the tag. I'd report it with the any-order check and the census on the five pinned apps. If that works for you, I'll claim it under "Any order".- added a commit that references this issue
on Oct 9, 2026 @eddygarcas, thanks. "Necessary, not sufficient" is the right framing, and the seven joins are exactly what the any-order brief was missing. They're now in it, linked back here, and the claim is recorded.
Points of contact for the PR:
- Overlap. A cloned type shares its children instead of copying them #656 changes how
Typayloads are shared, mostly mechanically across many files. The staged S2b and S2c commits touchunify_param_ty, the ivar merges and the harvest. Main comes first: whichever of A cloned type shares its children instead of copying them #656 and your PR lands second rebases, and the staged branch rebases onto both. - Running the checks. The canaries aren't on main until Opt-in canaries show whether the whole-program fixpoint settled #657 merges, so build the change on top of
fixpoint-nextto run the any-order check. There,PROBE_BASE=1gives main's behavior in the same binary, andRH_PRECISION_CENSUS=1adds the precision shares. - Item 4 may explain the unsound
@datacase in Discourse'sapp/jobs/base.rb(facts F8): an ivar seeded with{}, then written through[]=. It would make a good regression test. - Item 5. The staged sharing (the second half of S1) reuses a payload only when it is equal including record field order, because some targets emit fields in that order. Comparing rows regardless of order works as long as the stored form keeps source order, as you suggest.
- Item 7. The
strip_unknownpath, which turns a pending arm besidenilinto a gradualNil | untyped, looks like the missing link for the question to @dai199. It's noted in the soundness brief.
- Overlap. A cloned type shares its children instead of copying them #656 changes how
@bunnykong On
fixpoint-onemergeit's dropped in the join itself, not in the old harvest rules.- In
det.rs,norm_optremoves aVararm of a top-level union as ⊥ (L512). Below the top it becomespending_untyped; at the top nothing is kept. SoNil | Varnormalizes toNil. - Every harvested return goes through it under
RH_DET:lat_normon the first write (harvest_return.rs L197) andlat_joinindecide_harvested_return(L144-L148). The rules eddygarcas pointed at (stabilize_untyped_return_oscillationviastrip_unknown) aren't reached in the trial. On main they turn the same pair intoNil | untyped, notNil.
Should it be kept as unresolved? I think so, because a
Varbesidenilis often not pending. The body typer returns the sameVar(0)(unknown()) when dispatch finds nothing, and nothing refines it later:- a modeled class whose walk misses the method (send.rs L1585);
- a
TimeorDatemethod that isn't modeled (L1882-L1883); - a union whose arms all miss (L1921).
Reading those as ⊥ removes a real "unknown" and leaves
nil, which would explain receivers that arenilalone where the source builds an object. I haven't traced the 59 sites to confirm it.Until S2a's tag can tell the producers apart, a safe rule is to drop a
Vararm only while the slot can still change, and keep any that survives to quiescence as unresolved, counted as not fully typed. With the tag, those dispatch fallbacks would produce unresolved directly, and only the placeholder seeded before a slot is first computed would be ⊥.Separately, I'll take the
name_collisionfollow-up: parameter rows keyed by receiver side, the follow-up named in #551. It's a change on main, outside the staged branch; I'll report it here with the checks.- In
4 remaining items
@bunnykong Thanks for the trace. It also corrects my examples: the fallbacks are in the class walk and on unmodeled receivers, not the
Time/Datepath. Agreed that the precision has to come back by modeling those methods.@eddygarcas The row shape is settled: rows are keyed
(ClassId, Symbol, MethodReceiver)(ParamKeyinsrc/analyze/mod.rs), andapply_param_sitesand the concern fold still join withunify_param_typer row, so items 1 and 2 can build on that. The(class, name)viewApp::inferred_method_paramspublishes stays as it is too: the one review comment on it doesn't hold, because the controller's class-side helper clones are thehelper_methodinstance methods themselves, whose view calls feed the instance row. CI is green; I'll note here when #674 lands.- added a commit that references this issue
on Oct 9, 2026 @eddygarcas #674 is merged (c19524c). Two review fixes went in with it: the owner walk now looks for a
defon the call's side first, soChild.fetchreachesGrand.fetchpast aBase#fetch, and the S0 canaries' fingerprint reads rows by side (Class.name/Class#name). The row key isParamKey = (ClassId, Symbol, MethodReceiver)as described above.Items 3 to 6 are up as #705, against main. All emitted output is unchanged on all 12 targets, and there is a generated lattice-law harness and a regression test for the
@datashape. Onfixpoint-nextthe pooled fully-typed share moves 68.83% → 68.99% (S3) and 69.02% → 69.20% withPROBE_BASE=1.Items 1 and 2 are built on #674's rows, but one choice is yours before they go up. @bunnykong @dai199, once
unify_param_tytreats anuntypedarm inside a union like a bare one (item 2's associativity fix),Nil | untypedjoins toNil. On our runtime probe that typesRelation#order_key_of'srecordasNil, and +2 sites lose their type. It is the same nil-only case as S2b on Campfire. Three options:- (a) Only a non-nil concrete type absorbs
untyped, sounify(Nil, untyped)staysNil | untyped. This keeps the laws, never narrows to bareNil, and is my preference until S2a lands. - (b) Keep
untypedarms inside unions as they are. The result then still depends on grouping, so item 2 isn't really fixed. - (c) Hold items 1 and 2 until S2a's tag, then absorb only pending and unresolved values.
Would you rather have (a) now, or wait for (c)?
- (a) Only a non-nil concrete type absorbs
@eddygarcas, thanks for #705. (a) now, please. It keeps the laws and never narrows to bare
Nil.I'd keep (a)'s rule after S2a too, because (c) as written would likely bring the
nil-only receivers back. In theRH_DETtrial, 52 of its 59nil-only receivers came from permanent dispatch fallbacks, 49 of them registry gaps (28 missing methods, 21 unmodeled receivers), and none from a pending value (F18). Keeping every survivingVaras unresolved removed all 59, but fully typed fell to 64.01%, against 68.83% on S3. (a)'s rule is narrower: it keeps the arm only when no non-nil type sits beside it.So unresolved would stop being an identity: pending is the identity, unresolved is absorbed only by a non-nil concrete type, and gradual stays an arm.
start_new_session_for(user)still getsUser. If the tags share oneuntypedarm, gradual has to win when they merge, or grouping matters again.- added a commit that references this issue
on Oct 9, 2026 - added 9 commits that reference this issue
on Oct 9, 2026
Current state (updated Oct 10)
rubys gave the go-ahead to stage the work, and dai199 confirmed that S2 fits
harvest_return.rs; their suggestions shaped this plan. S2a to S3 are commits on the staged branch, with S2b to S3 behind flags that are off by default. A follow-up branch adds shuffled-order canaries, an error census and two fixes to the sharing. Milestone 2 has the staged measurements. The research is open in #669, where anyone can take on a problem, and records what came after, including the merged #674 (the receiver-side rows the proposal below calls a follow-up), #705 and #724.untypedexists: pending, gradual or unresolvedGates in roundhouse's terms, for S3 with every flag on against main
194f26cfat milestone 2 (first update, second); the October 10 re-measurement follows the table:checktimeOctober 10, on main
c210f226with #705 and #724 merged. The research corpus re-measured S3 on the five public apps. Each point names its fact in facts.md, and each fact links its receipts in the lab.Corrections to the original proposal
Ty::Bottom, which means "never returns" and prints asnever,!orbot. Pending values left at the end go through the entry-point policy, which settles them as gradualuntyped.Tyalready keepsVar,UntypedandBottomapart; the operations merge them (is_unknown,strip_unknown, the harvest fallback). S2a puts provenance onTy, which reverses Stabilize harvested returns that oscillate on Untyped/Var #505's choice to keep it off, and its commit message says why.untypedorVarat any depth, next to the share holdinguntypedand the runtime oracle's rejections. A count alone rewards rules that drop arms a value really has; the oracle catches those.Links: staged branch · follow-up branch · lab · research: #669 · related: #630 (class/instance dispatch, merged)
Original proposal (frozen)
Roundhouse's type analysis copies recursive types one level deeper each round, until the round cap or #584's size bound cuts them off. Stored as cycles instead, and computed by rules that never retract a conclusion when they learn more, those types stop growing, and the analysis settles on its own.
untypedA loop settles when a round changes no signature before the 12-round cap. Roundhouse runs three loops in turn: production; views and tests; and absorb, which re-runs production when views or tests moved one of its signatures.
Main fits in memory today because #584's bound cuts large carried types to
untyped. With each distinct piece stored once, no size bound is needed.The combined prototype is still slower on the large app: about 41 s, or 46–47 s with S1's sharing, against 31 s for the same base with a size bound. More of its expressions also contain
untypedsomewhere: 24.4% against 23.0%. On its own, the ordered worklist took about half the time (0.53× and 0.61× in two paired runs), though not yet with identical answers; bringing that speed into the combination is the next step.Baselines: the first two rows are main after #584 (
4be26816for the large app,07ce8255for the public apps); everything else isb28b17b6, from before #584, where the prototypes were built. In the proposal column, the memory rows are S1 alone, the chain row is S3 alone, and the rest are the combined prototype: S2 and S3 together, whose results S1's sharing doesn't change. Re-measurement on current main is under way. Each prototype is a patch in the lab, which reproduces every public number; the combined one isphase-c2.diff, research probes included. Conditions are listed at the end.How it works
Take the JSON-style normalizer from #518 (reproduction). Its true return type is recursive:
Each round re-types the method with the previous round's answer pasted in, so after k rounds the type is unfolded k levels deep. thomasklemm logged its parameter type growing 9, 15, 33, … 332k nodes until the run was killed at 4 GiB; this shape escapes #528's recursion cut. Since #584, a size bound stops the growth, and whatever it cuts becomes
untyped.Stored instead as a cycle, with a reference back to where the value came from (its origin), the same type is finite, and the loop can arrive at it.
Some rules, though, retract conclusions when they learn more:
untypedswallows the whole union, while parameter unification drops it.With rules like these, a loop can flip forever while nothing grows. As thomasklemm found in #518, such flips are what hold Mastodon and Discourse at their caps. When every rule is monotone, never retracting a conclusion when it learns more, and the loop starts from nothing known yet, it can only gain facts, and it stops, in any order of evaluation.
Each half fails without the other. Monotone rules without cycles keep climbing; cycles with rules that retract keep flipping. Together they settle on the least solution, the most precise answer the rules allow, in any order of evaluation. That is machine-checked in Lean 4 for a core calculus; the Rust prototype demonstrates settling, but not yet that its answer is the least one.
The approach isn't new. Research type checkers for unannotated Scheme programs used it in the 1990s to infer and print recursive types, and TypeProf takes a similar approach for Ruby; citations are in the credits. The staged plan also answers three concerns on record: an accumulating row contaminated by an early placeholder (68f4d828), the "larger typing-provenance change" #505 anticipated, and "a new
Tyvariant that every emitter handles" (#518).Proposed stages
How the stages land:
Each stage ships as small PRs; anything that changes results stays behind a flag until it passes the decision gates under Details. Emitting recursive types in generated code, new or raised caps, and inference state persisted between builds are out of scope.
Questions for the maintainers
untypedcan mean not computed yet, opted out or unresolved. Is splitting those three, with an explicit policy for entry points and for methods whose callers are never seen, acceptable as the "larger typing-provenance change" Stabilize harvested returns that oscillate on Untyped/Var #505 anticipated?Ty::Recthat emitters print asuntyped? The alternative, folding each slot's history where a round hands types to the next, needs noTychange, but it reads only history, so by finding 2 below it can't be exact on every acyclic chain; it is offered only as a fallback.Details
Why the loops don't settle: four findings
Growth (#518) and oscillation share a root: the loop has nothing to settle to. Oscillation is on record twice: rubys noted Mastodon methods that "flip between two types each round" (ccbafad6), and thomasklemm named it "the next convergence item" (#518). Today a cap, a bound (#584) and a stabilize rule (#521) make the loop stop anyway, and each hides the others' symptoms.
1. The memory is copies, not information (measured). A carried type is one that a round hands to the next: a method's return, a parameter row, an instance variable or a constant. On
b28b17b6with no size bound, the large app's largest carried type, a method return, has 24,052,607 tree nodes but only 127 distinct subtrees. The sampled signature state has 54.0M nodes and 17,663 distinct subtrees. As the largest type grows, its tree roughly triples each round, while its distinct subtrees grow by about nine.2. The answer is a cycle, not a finite tree (proved for the calculus). Each return, parameter row, instance variable and constant is a slot with one type for all calling contexts, so the rules are set constraints: each slot stands for a set of possible values, constrained by the code. With finitely many constructors and program sites, no negated constraints, and conditions restricted to sites that can build a value, their least solution is a regular tree language (Heintze–Jaffar): a set of trees that a finite grammar describes. On recursive data, that grammar has a cycle. A loop over finite trees holds depth k after round k, much as listing
x,xx,xxx, … one at a time never arrives atx*.The cycle must point to where a value came from, not to what it resembles. #528's cut and #584's bound see only a slot's own history, and a cutoff like that provably can't tell recursion from a long chain. Suppose it stops the recursive slot
X = Integer | Array[X]at round k, and what it cuts stays cut. On a chain of k methods, each wrapping the next one's result in anArray, the first slot has exactly the recursive slot's history for k rounds, so the cutoff fires there too and over-approximates forever. A recorded origin tells them apart: the chain has no back edge, no reference back to an earlier slot.3. Some rules aren't monotone, and some state is hidden (measured, with public reproductions). A rule is monotone if learning more about its inputs never retracts a conclusion about its output.
HashintoHash | Array[Integer], and the next turns it back intoHash: more input gives less output, so the row flips forever.untypedwins in dispatch, then is dropped. Dispatch on a union that containsuntypedreturns bareuntyped, discarding the other arms' results; parameter unification then drops thatuntyped. In the prototype's first configuration, this pair caused the last oscillation in Mastodon's absorb loop.untypedorVararms, Stabilize harvested returns that oscillate on Untyped/Var #521's stabilize keeps the first variant written. On Forem, global rounds and an ordered schedule both reach a state that one more full round leaves unchanged, yet 18 entries differ between them, each only in those arms: two different fixpoints of the same rules.4. Rounds repeat work (measured). On the large app, depending on the loop and round, 74–98% of method-body re-typings produce the same output as before. Only re-typings whose inputs didn't change can be skipped safely, so this is an upper bound on avoidable work. Long chains need many rounds: the cap went from 4 to 12 when lobsters needed 9 (be964432). In the measured visit order, a generated 64-method chain advances one method per round and exhausts the cap, despite #524's dirty set, which narrows production rounds to what the call graph says changed.
Why accumulation can work now, after two failed attempts
Accumulating, joining each round's result into what earlier rounds found instead of recomputing it, was tried twice, and the two attempts failed differently:
Untyped; round 2 seesStr; unify fuses them intoStr | Untyped, which isuntypedwith extra steps." An early placeholder had contaminated the row.Both point to the same lesson: accumulation is only safe when the analysis keeps the kinds of unknown apart.
Today one value carries two meanings that pull in opposite directions. Not computed yet (call it pending) is no evidence: it sits below every type, and the next round may replace it. The author opted out (call it gradual) is every value at once, and no round may remove it. Every rule that meets
untypedmust either keep it or drop it, and each choice is wrong for one meaning:Untypedis 68f4d828's contamination;untypedis how Stabilize harvested returns that oscillate on Untyped/Var #521's stabilize can strip a genuine arm, and a short public reproduction types anIntegerasString;S2 separates the meanings. Pending becomes ⊥ ("bottom"): joining ⊥ with any type
TgivesT. Gradual staysuntyped. Unresolved external or unsupported inputs become a third kind. With monotone rules and ⊥ as the starting value, every state the loop passes through lies below the final answer, so no round retracts what an earlier round concluded (the standard fixed-point theorems of Kleene and Knaster–Tarski). Operations that do narrow or retract, namely branch refinement, checks against declared types and retraction after source deletions, stay separate from this loop.Accumulation alone isn't enough. For a non-monotone rule F, the iteration X ↦ X ∪ F(X), which only ever adds facts, can stop at a state where F(X) ⊆ X but F(X) ≠ X: some facts survive only because an earlier round produced them. Leastness and order independence don't follow.
Pieces are already in place: 16e1ead2 makes
union_ofa law-abiding join, with exhaustive tests of its laws, and #565 makes pending parameter joins commutative. Neither certifies the whole analyzer. #505 called the split "a larger typing-provenance change", and it is one: S2 needs the split at the producers, plus an explicit policy for entry points and unseen callers.Stages in detail
S0: canaries and counters (no behaviour change). Part of this already exists:
fixpoint_rounds();S0 adds:
S1: shared types with DAG-aware operations (intended to preserve behaviour). ccbafad6 (rubys) shared the constant registry instead of copying it and took Mastodon's
checkfrom 12.9 to 4.6 s; #171 (tobi) shared constant scopes. S1 extends sharing to the types themselves.Storing the 127 distinct subtrees once isn't enough on its own: comparisons, hashes, substitutions and rewrites still walk the 24-million-node tree they stand for. Two rules fix that:
Measured on
b28b17b6with no size bound; #528's cut and the round caps still apply:untyped-count ceiling tests hold at 0, 304 and 523.The first step of the change touches 108 files. S1 hasn't been measured together with #584's bound; that measurement comes first.
S2: monotone rules from ⊥, landed together with recursive types by origin.
Recursive references. Inside a dependency component, a type refers to another slot's type by a (slot, program position) pair instead of embedding a copy. This is the classic set-based technique of one variable per program site; TypeProf v2 does something similar for container element types (related work).
Tygains an analysis-onlyTy::Rec, which takes nine mechanical exhaustive-match updates, including fallbacks in eight emitter type renderers and the IDE. Emitters printuntypedwhere a reference remains. That narrows thomasklemm's concern, "a newTyvariant that every emitter handles", to small fallback changes now, with typed recursive emission as later work.A fallback needs no
Tychange: where a round hands its types to the next, fold each slot's history into a recursive type kept in a side table, and let the typer see only a depth-limited unfolding of it, withuntypedbelow. It reads only the slot's history, so by finding 2 it can't be exact on every acyclic chain; it is offered only as a fallback.Monotone rules from ⊥:
Why the halves land together (measured on the five public apps unless noted):
untypedfrom swallowing the other arms in dispatch grows Discourse's final state 28× (1.09M → 30.9M nodes, 5.1 GiB). Removing one non-monotone rule can uncover growth that the old oscillation was holding in check.07ce8255, which includes A type the fixpoint carries between rounds stays bounded (#518) #584, leaves Mastodon and Discourse at their caps and adds 32 errors on Mastodon. On the large app, A type the fixpoint carries between rounds stays bounded (#518) #584's bound fires 5,243 times instead of 2,036.S2's earlier prototype,
fold2, without the ordered worklist, on the large app: 3.58 GiB and 78.2 s, against 34.5 s for the historical bounded control. The tests and absorb loops settle; production still runs to the cap.S3: ordered rounds, after S2.
The change:
Read sets must include every semantic input; a changed input may still produce unchanged output.
Why after S2. The least-solution theorem for dependency-ordered solving assumes a monotone system with complete dependencies, so it applies once S2 holds.
In Spinel:
Measured alone (S3 without S1 or S2), on
b28b17b6with the historical bounded control's size bound:The answers differ, so this is not a like-for-like speedup. The scheduler freezes 139 units (the bodies it types) after 12 changes each within a phase, an extra full round still moves 68 slots after production and 118 after absorb, and errors rise by 8. Inside the combined prototype, no unit freezes. On the large app, one dependency component holds 17,284 of 21,495 units (80%), so the gain comes from selective re-typing within components, not from small components.
S4: precision and printing.
Per-argument-shape summaries (modeled). Today a method gets one summary for all its calls. In a finite model of the normalizer, typing the body separately for Hash, Array and other arguments (three body rows instead of one) removes the 8 spurious receiver arms at its two
mergecalls. The complete shell, the smallest refinement that makes the analysis exact for these operations (Giacobazzi–Ranzato–Scozzari), shows that two summaries suffice for those calls: one for Hash arguments and one for everything else. Splitting has a cost: Endoh's commit e68384616 removed union-argument expansion, turning a 27-overload example into one union signature. So the number of summaries per method must be measured. Separate class-side and instance-side parameter rows remain a follow-up named in #551.Printing. Before export, resolve alias cycles that pass through no type constructor (
type a = b | nilwithtype b = a), then validate withrbs; version 3.10.0 accepts the JSON alias from How it works. Roundhouse's reader currently rejects multiple overloads and leaves cyclic aliases unread (src/rbs.rs), so the reader comes first. The Rust-target gaps #589 records for methods that walk recursive data,serde_json::Valuelackingtransform_valuesandcase value when Hash / when Arrayrendering as a type-blindmatch, need the operations and branch tests lowered too, not just a type.What the combined prototype shows, and what it doesn't
The combined prototype is S2's recursive references and monotone rules with S3's worklist, on
b28b17b6, with no size bound; adding S1's sharing changes no result, only memory and time. Its whole diff, sharing included, isphase-c2.diff, run with the flags ofphase-c2ain configurations.json.Large app. Each configuration has two timing runs, back to back with no other benchmark running; separate instrumented runs check the result, three without sharing and one with it.
Public apps:
Three integration fixes made the difference:
The evidence supports three claims, of decreasing strength:
untyped), because order-dependent rules remain.sort.to_h) aren't propagated. Traces from the other 12 small reproductions are accepted.Precision. On the large app, 24.39% of the 1,232,382 expressions that carry a type contain
untypedat some depth in the combined prototype, against 23.02% in the historical bounded control; on the five public apps, 14.45% against 14.18%. Counted by cause, most of the remaininguntypedcomes from parameters that no caller gives a type, from failed dispatch, and from pending placeholders that are never replaced; theuntypedprinted where a reference remains doesn't explain the gap on its own.What is mechanized (Lean 4)
The fixed-point results are classical textbook theorems: Kleene and Knaster–Tarski fixed points, set-based analysis, and chaotic iteration. Several have been mechanized before, in Isabelle and Coq. The new part is a Lean 4 proof for a small model of the analysis with finitely many program sites, plus a proof that this model agrees with a relational statement of its rules (
Equiv.pt_iff). Nothing proves that the Rust analyzer implements the model.The build has no unproved steps (
sorry,admit), no trusted evaluation (native_decide) and only the standard axioms. It proves:Counterexamples show why the hypotheses can't be dropped:
Var, reaches a fixpoint that isn't least;A starting state proven to lie below the answer may replace ⊥, and derived caches are allowed with correct invalidation.
Not proposed, decision gates and ownership
Not proposed:
untypedwhere a reference remains);untypedused to hide are reported, and counted separately).Decision gates:
untypedback at main's level.Ownership:
harvest_return.rs(A recursive method's harvested return stops unrolling; a gap cause is the same on every run #528, A call through x.class does not feed the instance method of that name #551; dai199) and the parameter join (Make a module parameter's call-site join order-independent (#209) #565; eddygarcas).dirty_retype.rsand the scheduling thomasklemm deferred in Stabilize harvested returns that oscillate on Untyped/Var #521 and began in Narrow production typing rounds with a call-graph dirty set #524.Each stage is offered as small PRs against that work, not around it.
Conditions and reproduction
Bases:
b28b17b6unless a commit is named.b28b17b6with an experimental depth-8, 256-node bound. It is not A type the fixpoint carries between rounds stays bounded (#518) #584, whose depth-16, 512-node bound applies where types are stored.4be26816for the large app and07ce8255for the public apps; live main has moved on since.Evidence labels:
Inputs:
The lab: bunnykong/roundhouse-fixpoint-lab holds the recipes, manifests, patches with their flags, runtime oracle and Lean project. Each result names its base, flags and command. Wall times depend on the host.
Credits
Upstream:
x.classfiltering;roundhouse checkfast on large apps #171's shared constant scopes and Meta PR 2: fixes from running roundhouse over a large internal Rails codebase #503's fixes from a large Rails codebase;Precedent:
The full list, with links, is in RELATED-WORK.md.