Repository navigation
Conversation
Allocate code chunks from a single reserved region aligned to 512MB, instead of wherever mmap finds room. In a process whose address space is already fragmented, such as a large application that loads Scheme as a library, code chunks could land gigabytes apart, and on Apple arm64 calls and returns whose targets lie in a different 1GB- or 4GB-aligned region than the branch mispredict far more often. The same program then ran up to about 1.7x slower for the life of the process, depending on where its code happened to land. The region is reserved PROT_NONE and made accessible piece by piece. Freed chunks release their pages with madvise and go on a free list for reuse; they stay accessible, since macOS refuses to change the protection of MAP_JIT memory once it has been made executable. If the reservation fails or the region fills, code chunks fall back to plain mmap. Other platforms are unchanged.
|
I ran the benchmark on my MacBook Pro with 32 GB of RAM on an Apple M2 Pro: % scheme --script repro.ss |
|
Thanks for running it. Could you paste the last line If it says |
Chez places each code chunk wherever mmap finds room. Inside a host process whose address space is already crowded, the chunks can land gigabytes apart, and on Apple arm64 calls and returns that cross a 1 GB or 4 GB-aligned boundary mispredict far more often, so a library ran up to 1.7x slower for the life of such a process (#1246). The library stub now includes host/chez/stub/jolt_code_region.h, which on macOS arm64 defines hidden mmap/munmap in the library image. The statically linked Chez kernel binds to them, and they carve anonymous MAP_JIT mappings (code chunks) out of one 512 MB region aligned to its own size, reusing freed ranges. Everything else, and anything once the region is full, goes to the real calls. Hidden, so the host process is untouched. The same fix for Chez itself is cisco/ChezScheme#1074; once jolt builds against a Chez that has it, this header can go. The build-lib smoke gains a code_churn export, which compiles and drops code so chunks are freed and reused, and on macOS arm64 a check that the library defines mmap locally without exporting it.
|
Thanks, that line settles it, and not in the PR's favour: your There were two problems. Collections kept freeing segments in the chunks below the boundary, and new code went back into them, so whether the fill held depended on when a collection ran. And even when the benchmark's code did start out past the boundary, a collection during the timed run could copy it back below. My earlier 234-244 ms numbers came from runs like that, so they overstate the effect. I've updated the scripts in the description. Automatic collection is now off from the start until the benchmark is compiled. Then the benchmark's code (every procedure in On an M1 Pro (macOS 26), 10 interleaved runs per cell:
With this PR, So the effect is about 1.36x on the M1 Pro, not 2.4x. Could you run the new |
|
|
Thanks for rerunning it. So the M2 Pro doesn't care: 97 ms with the code past the boundary, same as the control, while the M1 Pro pays about 1.36x. It looks like an M1-generation quirk. I've narrowed the description, release note and code comment to say that. The region still helps M1 machines and costs nothing on later ones, but if you'd rather not carry it for M1 alone, I'm fine closing this. |
|
Thanks for the PR! I'm reluctant to merge mainly because I'm not sure of all the implications of reserving half a gigabyte of memory. The big |
On Apple arm64, allocate code chunks from one reserved region aligned to 512MB, so all code stays inside a single aligned window of the address space. Today each code chunk goes wherever
mmapfinds room. In a process whose address space is already fragmented, typically a large application that loads Chez Scheme as a library, code chunks can land gigabytes apart, and on an M1 the same program then runs up to about 1.7x slower for the life of that process. An M2 Pro shows no such penalty (see the comments).What happens
On an M1 Pro (macOS 26), the slowdown follows where code chunks sit relative to aligned address boundaries:
Scheme returns (
ldur x10, [x20]; br x10) and most calls are indirect branches, and when a hot branch and its target lie in different 1GB- or 4GB-aligned regions, they mispredict far more often. With a modified kernel that places code chunks deliberately:So it is not distance as such. It is crossing a 1GB or 4GB boundary.
Reproducing it with plain Chez Scheme
Revised: an earlier version of these scripts could time a run whose code had not crossed the boundary, or had been moved back by a collection mid-run, and its numbers overstated the effect. See the comments.
repro.sscompiles a call-heavy benchmark (bench.ss, a small boids-style step over lists of flonum vectors) after start-up, into a fresh code chunk. Withfill, it first takes all free address space below the next 4GB boundary above the run-time system's code, so that chunk has to be placed past the boundary. That is the only difference between the two modes. Automatic collection is off until the benchmark is compiled, the benchmark's code is then locked so it can't move during the timed run, and the last line says which case the time is for.M1 Pro (macOS 26), 10 interleaved runs per cell:
fillmain(35c2f4b3)main+ this PRWith this change,
fillcan't push the benchmark's code past the boundary, because code chunks come from the reserved region, so both modes run at the control's speed. One of those ten runs took 175 ms; the other nine were 109-116 ms. Homebrew 10.4.1 behaves likemain: 114-116 ms against 152-155 ms in 3 pairs.repro.ss
bench.ss
The effect was found with a Chez-based library loaded into a process that had already reserved 3GB of address space. Without this change, 10 of 24 such processes ran 10% to 75% slow (5.43 to 9.43 ms per step). With it, none of 24 did (5.38 to 5.59 ms).
The change
c/version.h: on Apple arm64 withoutWRITE_XOR_EXECUTE_CODE(the case that usesMAP_JIT), defineS_CODE_REGION_BYTESas 512MB.c/segment.c: whenS_CODE_REGION_BYTESis defined, the first code-chunk allocation reserves twice the region sizePROT_NONE | MAP_JIT, keeps an aligned window of the region size and unmaps the rest. A window aligned to its own size never straddles a boundary larger than itself. Code chunks are carved from the window, made accessible withmprotectthe first time a range is handed out. A freed chunk's pages are released withmadvise(MADV_FREE)and its range goes on a sorted, coalescing free list for reuse. It stays accessible, because macOS refuses (EACCES) to change the protection ofMAP_JITmemory once it has been made executable. If the reservation fails or the region is full, code chunks fall back to plainmmap, as before.Chunks are allocated with the allocation mutex held and freed by the collector after sweeping, so the region needs no lock of its own. Region memory is only handed out for
zerofill == 0requests, since a reused range is not zero-filled, and code chunks are always requested that way.Other platforms are unchanged. The penalty was measured on an M1 Pro; on an M2 Pro the same reproduction runs at the same speed whether or not the code crosses the boundary (thanks @burgerrg). On later cores the region costs nothing, since code just lands in one place.
Testing
make test-some-fastontarm64osx(M1 Pro, macOS 26) passes in all three configurations, andmisc.mo, with the new mat, passes on three separate runs.code-placementinmats/misc.ms(arm64osx and tarm64osx only): code allocated after compiling 2000 procedures and a full collection is in the same 512MB-aligned window as code allocated before, and code ranges freed by a collection get reused and still run. It compares instruction addresses fromforeign-callable-entry-point. Stock Chez Scheme also passes it in an unfragmented process, where code chunks happen to stay close anyway, so it guards the placement and reuse paths rather than demonstrating the slowdown.repro.ssdemonstrates the slowdown.