Skip to content

keep code in one aligned region on Apple arm64 - #1074

Open
burinc wants to merge 2 commits into
cisco:mainfrom
burinc:arm64-code-region
Open

burinc wants to merge 2 commits into
cisco:mainfrom
burinc:arm64-code-region

Conversation

@burinc

@burinc burinc commented Oct 5, 2026 •

Copy link
Copy Markdown

On Apple arm64, allocate code chunks from one reserved region aligned to 512MB, so all code stays inside a single aligned window of the address space. Today each code chunk goes wherever mmap finds room. In a process whose address space is already fragmented, typically a large application that loads Chez Scheme as a library, code chunks can land gigabytes apart, and on an M1 the same program then runs up to about 1.7x slower for the life of that process. An M2 Pro shows no such penalty (see the comments).

What happens

On an M1 Pro (macOS 26), the slowdown follows where code chunks sit relative to aligned address boundaries:

  • Instructions retired are identical in fast and slow processes (66.91e9 to 66.93e9 for one workload), while cycles go from 9.9e9 to 16.9e9.
  • Instruments' CPU Counters ("CPU Bottlenecks" mode) show the share of cycles spent on discarded (mispredicted) work rising from about 10% in a fast process to 21% in a 1.15x-slow one and 40% in a 1.75x-slow one, while the useful share falls from 85% to 73% and 49%, in proportion to the slowdown.
  • Sampling puts the extra time on procedure entries and return points.

Scheme returns (ldur x10, [x20]; br x10) and most calls are indirect branches, and when a hot branch and its target lie in different 1GB- or 4GB-aligned regions, they mispredict far more often. With a modified kernel that places code chunks deliberately:

code chunks time per step
adjacent, or up to 192MB apart 5.4 ms
256MB to 1GB apart 5.9 to 6.2 ms
2GB to 4GB apart 9.5 to 9.7 ms
adjacent, straddling a 256MB boundary 5.4 ms
adjacent, straddling a 1GB boundary 5.9 ms
adjacent, straddling a 4GB boundary 8.0 to 8.1 ms

So it is not distance as such. It is crossing a 1GB or 4GB boundary.

Reproducing it with plain Chez Scheme

Revised: an earlier version of these scripts could time a run whose code had not crossed the boundary, or had been moved back by a collection mid-run, and its numbers overstated the effect. See the comments.

repro.ss compiles a call-heavy benchmark (bench.ss, a small boids-style step over lists of flonum vectors) after start-up, into a fresh code chunk. With fill, it first takes all free address space below the next 4GB boundary above the run-time system's code, so that chunk has to be placed past the boundary. That is the only difference between the two modes. Automatic collection is off until the benchmark is compiled, the benchmark's code is then locked so it can't move during the timed run, and the last line says which case the time is for.

scheme --script repro.ss         # benchmark code next to the run-time code
scheme --script repro.ss fill    # benchmark code past the next 4GB boundary

M1 Pro (macOS 26), 10 interleaved runs per cell:

control fill
main (35c2f4b3) 109-114 ms, median 111 144-154 ms, median 151, past the boundary in all 10
main + this PR 108-113 ms, median 111 109-175 ms, median 112, code never crosses

With this change, fill can't push the benchmark's code past the boundary, because code chunks come from the reserved region, so both modes run at the control's speed. One of those ten runs took 175 ms; the other nine were 109-116 ms. Homebrew 10.4.1 behaves like main: 114-116 ms against 152-155 ms in 3 pairs.

repro.ss
;; Usage: scheme --script repro.ss [fill]
;;
;; Times bench.ss, a call-heavy benchmark compiled after start-up, and says
;; where its code was placed relative to the run-time system's code.
;;
;; Both modes allocate enough small locked code objects (foreign callables) to
;; use up the existing code chunks and start a fresh one, then compile the
;; benchmark, whose code goes into that fresh chunk. With `fill`, the script first takes all free address space
;; below the next 4GB boundary above the run-time system's code, so the fresh
;; chunk has to be placed past that boundary. That is the only difference
;; between the two modes.
;;
;; Automatic collection is off from the start until the benchmark is compiled:
;; a collection frees segments in the chunks below the boundary, and new code
;; would go there instead. The benchmark's code is then locked, so collections
;; during the timed run don't copy it elsewhere. The script prints where it is
;; before and after the run, and a last line saying which case the time is for.
(load-shared-object "libc.dylib")
(define mmap (foreign-procedure "mmap" (uptr size_t int int int long) uptr))
(define munmap (foreign-procedure "munmap" (uptr size_t) int))
(define (code-of proc) (#%$object-address (#%$closure-code proc) 0))
(define (next-4gb a) (* (+ (quotient a (expt 2 32)) 1) (expt 2 32)))

(define (fill-below! limit)
  ;; take the lowest free range again and again, largest pieces first, until
  ;; the kernel hands back an address past `limit` (PROT_NONE, MAP_PRIVATE|MAP_ANON)
  (for-each
   (lambda (size)
     (let loop ()
       (let ([p (mmap 0 size 0 #x1002 -1 0)])
         (if (< p limit) (loop) (munmap p size)))))
   (list (expt 2 28) (expt 2 24) (expt 2 21) (expt 2 18) (expt 2 14))))

(define callables '())
(define (callable-address)
  (let ([c (foreign-callable (lambda () 0) () int)])
    (lock-object c)
    (set! callables (cons c callables))
    (foreign-callable-entry-point c)))

(define (use-up-code-space!)
  ;; 20000 small code objects, about three times what it takes to fill the
  ;; code chunks there are and start a fresh one; with `fill` that fresh chunk
  ;; has nowhere to go but past the boundary. Both modes allocate the same
  ;; number, so collections during the run have the same locked objects to see.
  (do ([n 0 (+ n 1)]) ((= n 20000)) (callable-address)))

(define (benchmark-code-objects)
  ;; the code of every procedure bench.ss defines, and of the lambdas inside
  ;; them (reached through relocations), but not the run-time system's, which
  ;; is in the static generation and never moves
  (let loop ([todo (map (lambda (name) (#%$closure-code (top-level-value name)))
                        '(v+ v- v* vlen vmean steer step new-flock run-benchmark))]
             [seen '()])
    (cond
      [(null? todo) seen]
      [(or (memq (car todo) seen)
           (> (#%$generation (car todo)) (collect-maximum-generation)))
       (loop (cdr todo) seen)]
      [else
       (let ([c (car todo)])
         (loop (append (filter #%$code? (((inspect/object c) 'reloc) 'value)) (cdr todo))
               (cons c seen)))])))

(define fill? (and (member "fill" (command-line-arguments)) #t))
(define runtime-code (code-of map))
(define boundary (next-4gb runtime-code))

(let ([handler (collect-request-handler)])
  (collect-request-handler void)
  (when fill? (fill-below! boundary))
  (use-up-code-space!)
  (load "bench.ss")
  (for-each lock-object (benchmark-code-objects))
  (collect-request-handler handler))

(define (benchmark-past?)
  (let* ([bench-code (code-of step)]
         [past? (> bench-code boundary)])
    (printf "run-time code ~x, benchmark code ~x (~a the 4GB boundary ~x)\n"
            runtime-code bench-code (if past? "past" "below") boundary)
    past?))

(let* ([before (benchmark-past?)]
       [_ (run-benchmark)]
       [after (benchmark-past?)])
  (printf
   (cond
     [(not (eq? before after))
      "a collection moved the benchmark's code during the run, so the time above mixes both cases\n"]
     [(and fill? after)
      "the benchmark's code ran past the boundary: the case being measured\n"]
     [fill?
      "the fill did not push the benchmark's code past the boundary, so this run is not that case (expected when code is kept in one region)\n"]
     [after
      "the benchmark's code is past the boundary even without `fill`\n"]
     [else
      "the benchmark's code ran next to the run-time code: the control\n"])))
bench.ss
;; A boids-style step over 100 birds, as lists of flonum vectors: lots of
;; small non-tail calls into closures and the run-time system.
(define (v+ a b) (vector (fl+ (vector-ref a 0) (vector-ref b 0)) (fl+ (vector-ref a 1) (vector-ref b 1)) (fl+ (vector-ref a 2) (vector-ref b 2))))
(define (v- a b) (vector (fl- (vector-ref a 0) (vector-ref b 0)) (fl- (vector-ref a 1) (vector-ref b 1)) (fl- (vector-ref a 2) (vector-ref b 2))))
(define (v* a k) (vector (fl* (vector-ref a 0) k) (fl* (vector-ref a 1) k) (fl* (vector-ref a 2) k)))
(define (vlen a) (flsqrt (fl+ (fl* (vector-ref a 0) (vector-ref a 0)) (fl* (vector-ref a 1) (vector-ref a 1)) (fl* (vector-ref a 2) (vector-ref a 2)))))
(define zero (vector 0.0 0.0 0.0))
(define (vmean vs) (v* (fold-left v+ zero vs) (fl/ 1.0 (fixnum->flonum (length vs)))))
(define (steer b flock)
  (let* ([p (car b)]
         [near (filter (lambda (o) (let ([d (vlen (v- (car o) p))]) (and (fl> d 0.0) (fl< d 6.0)))) flock)]
         [close (filter (lambda (o) (fl< (vlen (v- (car o) p)) 1.5)) near)]
         [sep (fold-left v+ zero (map (lambda (o) (v- p (car o))) close))]
         [ali (if (null? near) zero (v- (vmean (map cdr near)) (cdr b)))]
         [coh (if (null? near) zero (v- (vmean (map car near)) p))])
    (v+ (v+ (v* sep 1.6) (v* ali 0.8)) (v* coh 1.2))))
(define (step flock dt)
  (map (lambda (b) (let ([v (v+ (cdr b) (v* (steer b flock) dt))]) (cons (v+ (car b) (v* v dt)) v))) flock))
(define (new-flock n)
  (let loop ([i 0] [acc '()])
    (if (= i n) acc
        (loop (+ i 1) (cons (cons (vector (fl* 10.0 (sin (fixnum->flonum i))) (fl* 10.0 (cos (fixnum->flonum i))) (fl* 0.1 (fixnum->flonum i)))
                                  (vector 1.0 0.5 0.25)) acc)))))
(define (run-benchmark)
  (let ([flock (new-flock 100)])
    (do ([i 0 (+ i 1)]) ((= i 30)) (set! flock (step flock (/ 1.0 60))))
    (let ([t0 (real-time)])
      (do ([i 0 (+ i 1)]) ((= i 300)) (set! flock (step flock (/ 1.0 60))))
      (printf "~a ms for 300 steps\n" (- (real-time) t0)))))

The effect was found with a Chez-based library loaded into a process that had already reserved 3GB of address space. Without this change, 10 of 24 such processes ran 10% to 75% slow (5.43 to 9.43 ms per step). With it, none of 24 did (5.38 to 5.59 ms).

The change

  • c/version.h: on Apple arm64 without WRITE_XOR_EXECUTE_CODE (the case that uses MAP_JIT), define S_CODE_REGION_BYTES as 512MB.
  • c/segment.c: when S_CODE_REGION_BYTES is defined, the first code-chunk allocation reserves twice the region size PROT_NONE | MAP_JIT, keeps an aligned window of the region size and unmaps the rest. A window aligned to its own size never straddles a boundary larger than itself. Code chunks are carved from the window, made accessible with mprotect the first time a range is handed out. A freed chunk's pages are released with madvise(MADV_FREE) and its range goes on a sorted, coalescing free list for reuse. It stays accessible, because macOS refuses (EACCES) to change the protection of MAP_JIT memory once it has been made executable. If the reservation fails or the region is full, code chunks fall back to plain mmap, as before.

Chunks are allocated with the allocation mutex held and freed by the collector after sweeping, so the region needs no lock of its own. Region memory is only handed out for zerofill == 0 requests, since a reused range is not zero-filled, and code chunks are always requested that way.

Other platforms are unchanged. The penalty was measured on an M1 Pro; on an M2 Pro the same reproduction runs at the same speed whether or not the code crosses the boundary (thanks @burgerrg). On later cores the region costs nothing, since code just lands in one place.

Testing

  • make test-some-fast on tarm64osx (M1 Pro, macOS 26) passes in all three configurations, and misc.mo, with the new mat, passes on three separate runs.
  • New mat code-placement in mats/misc.ms (arm64osx and tarm64osx only): code allocated after compiling 2000 procedures and a full collection is in the same 512MB-aligned window as code allocated before, and code ranges freed by a collection get reused and still run. It compares instruction addresses from foreign-callable-entry-point. Stock Chez Scheme also passes it in an unfragmented process, where code chunks happen to stay close anyway, so it guards the placement and reuse paths rather than demonstrating the slowdown. repro.ss demonstrates the slowdown.
  • The plain-Chez reproduction above, before and after.

Allocate code chunks from a single reserved region aligned to 512MB,
instead of wherever mmap finds room. In a process whose address space
is already fragmented, such as a large application that loads Scheme
as a library, code chunks could land gigabytes apart, and on Apple
arm64 calls and returns whose targets lie in a different 1GB- or
4GB-aligned region than the branch mispredict far more often. The
same program then ran up to about 1.7x slower for the life of the
process, depending on where its code happened to land.

The region is reserved PROT_NONE and made accessible piece by piece.
Freed chunks release their pages with madvise and go on a free list
for reuse; they stay accessible, since macOS refuses to change the
protection of MAP_JIT memory once it has been made executable. If the
reservation fails or the region fills, code chunks fall back to plain
mmap. Other platforms are unchanged.
@burgerrg

burgerrg commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

I ran the benchmark on my MacBook Pro with 32 GB of RAM on an Apple M2 Pro:

% scheme --script repro.ss
92 ms for 300 steps
run-time code 106DE1BBF, benchmark code 107E1D34F (below the 4GB boundary 200000000)
% scheme --script repro.ss fill
98 ms for 300 steps
run-time code 109B5DBBF, benchmark code 10AB9134F (below the 4GB boundary 200000000)

@burinc

burinc commented Oct 6, 2026

Copy link
Copy Markdown
Author

Thanks for running it. Could you paste the last line repro.ss prints in the fill run? It says where the benchmark's code ended up (past or below the boundary), so it shows whether the fill actually moved it.

If it says past, then the M2 Pro pays about 7% where the M1 Pro pays 2.4x, which would mean the M2's branch predictor handles targets in another 4GB region much better. I only have an M1 Pro, so I can't check that myself. In that case the change matters mostly on M1-generation cores, and it costs nothing on later ones, since code just lands in one region. I'd update the description and the release-notes entry to say so, rather than claim a slowdown on Apple arm64 in general.

yogthos pushed a commit to jolt-lang/jolt that referenced this pull request Oct 6, 2026
Chez places each code chunk wherever mmap finds room. Inside a host
process whose address space is already crowded, the chunks can land
gigabytes apart, and on Apple arm64 calls and returns that cross a
1 GB or 4 GB-aligned boundary mispredict far more often, so a library
ran up to 1.7x slower for the life of such a process (#1246).

The library stub now includes host/chez/stub/jolt_code_region.h, which
on macOS arm64 defines hidden mmap/munmap in the library image. The
statically linked Chez kernel binds to them, and they carve anonymous
MAP_JIT mappings (code chunks) out of one 512 MB region aligned to its
own size, reusing freed ranges. Everything else, and anything once the
region is full, goes to the real calls. Hidden, so the host process is
untouched. The same fix for Chez itself is cisco/ChezScheme#1074; once
jolt builds against a Chez that has it, this header can go.

The build-lib smoke gains a code_churn export, which compiles and drops
code so chunks are freed and reused, and on macOS arm64 a check that
the library defines mmap locally without exporting it.
@burinc

burinc commented Oct 6, 2026

Copy link
Copy Markdown
Author

Thanks, that line settles it, and not in the PR's favour: your fill run left the benchmark code below the boundary (10AB9134F), so it ran the same case as the control, and the script printed a time anyway. That's a bug in repro.ss, not something the M2 showed.

There were two problems. Collections kept freeing segments in the chunks below the boundary, and new code went back into them, so whether the fill held depended on when a collection ran. And even when the benchmark's code did start out past the boundary, a collection during the timed run could copy it back below. My earlier 234-244 ms numbers came from runs like that, so they overstate the effect.

I've updated the scripts in the description. Automatic collection is now off from the start until the benchmark is compiled. Then the benchmark's code (every procedure in bench.ss and the lambdas inside them) is locked, so it can't move during the timed run. The script prints where the code is before and after the run, and a last line saying which case the time is for: the control, the case being measured, or a fill that didn't get the code past the boundary.

On an M1 Pro (macOS 26), 10 interleaved runs per cell:

control fill
main (35c2f4b3) 109-114 ms, median 111 144-154 ms, median 151, past the boundary in all 10
main + this PR 108-113 ms, median 111 109-175 ms, median 112, code never crosses

With this PR, fill can't push the code past the boundary, because code chunks come from the reserved region. The script says so and times it anyway. One of those ten runs took 175 ms; the other nine were 109-116 ms.

So the effect is about 1.36x on the M1 Pro, not 2.4x. Could you run the new repro.ss and bench.ss from the description on your M2 Pro? If the last line of the fill run says "past the boundary" and the time matches the control, the M2's predictor handles it, and I'll narrow the PR to say so.

@burgerrg

burgerrg commented Oct 7, 2026

Copy link
Copy Markdown
Contributor
% scheme --script repro.ss
run-time code 1073E1BBF, benchmark code 105F05BCF (below the 4GB boundary 200000000)
97 ms for 300 steps
run-time code 1073E1BBF, benchmark code 105F05BCF (below the 4GB boundary 200000000)
the benchmark's code ran next to the run-time code: the control
% scheme --script repro.ss fill
run-time code 10B92DBBF, benchmark code 30066DBCF (past the 4GB boundary 200000000)
97 ms for 300 steps
run-time code 10B92DBBF, benchmark code 30066DBCF (past the 4GB boundary 200000000)
the benchmark's code ran past the boundary: the case being measured

@burinc

burinc commented Oct 8, 2026

Copy link
Copy Markdown
Author

Thanks for rerunning it. So the M2 Pro doesn't care: 97 ms with the code past the boundary, same as the control, while the M1 Pro pays about 1.36x. It looks like an M1-generation quirk.

I've narrowed the description, release note and code comment to say that. The region still helps M1 machines and costs nothing on later ones, but if you'd rather not carry it for M1 alone, I'm fine closing this.

@mflatt

mflatt commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Thanks for the PR! I'm reluctant to merge mainly because I'm not sure of all the implications of reserving half a gigabyte of memory. The big mmap might be rejected by a per-process limit on memory use, for example, keeping a program that would otherwise run in less memory from proceeding — or causing it to fail shortly after this allocation, even if the immediate mmap succeeds. There are ways to cover those bases, no doubt, but I would lean toward the simplicity of not adding this in the main branch.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants