DO NOT SUBMIT: hugepage follow-up ideas as of 2026-09-12

Alpha has no leaf entry above the PTE level. A PMD entry covers 8MB and
the largest granularity hint (GH) block is 4MB, so anything in core mm
that assumes PMD or PUD leaves does not fit. Everything below has to work
the way arm64's contiguous PTEs do: several PTEs describing one physical
block.

Worth building
==============

1. Transparent GH for large folios (biggest win, most work)

   Model: arch/arm64/mm/contpte.c. When set_ptes() maps an aligned 64K,
   512K or 4MB folio, set the hint on the whole block, and split it back
   into single pages on mprotect, partial munmap, CoW and so on.

   Payoff: ordinary programs get the TLB benefit without hugetlbfs. That
   covers anon mTHP, large page cache folios (ext4, xfs, tmpfs) and exec
   text read ahead at 64K through exec_folio_order() (mm/filemap.c).
   The EV7 measurement (555 vs 301 cycles per access, base pages vs 4MB)
   is the justification.

   Blocker: every one of those paths depends on
   CONFIG_TRANSPARENT_HUGEPAGE, which needs HAVE_ARCH_TRANSPARENT_HUGEPAGE,
   which assumes PMD leaves.
   - Without THP, mapping_max_folio_size_supported() returns PAGE_SIZE
     (include/linux/pagemap.h).
   - The runtime off switch does not help: thp_disabled_by_hw() also
     disables mTHP (mm/memory.c).
   So this needs a core mm change: let an arch enable THP with the PMD
   order never allowed, and stub out pmd_trans_huge() and friends. That
   is an upstream discussion, not just arch code.

   The rules that every PTE in a block must match in bits <15:0>, and
   that young/dirty apply to the whole block, are already implemented in
   arch/alpha/mm/hugetlbpage.c. That code can be moved somewhere shared.

2. Huge vmalloc through GH (huge-vmap is TODO for alpha)

   arm64 does this with contiguous PTEs: arch_vmap_pte_supported_shift()
   and arch_vmap_pte_range_map_size(). It needs HAVE_ARCH_HUGE_VMAP plus
   stub pmd_set_huge()/pud_set_huge() that return 0.

   The benefit is small. KSEG already maps most kernel memory with no TLB
   entries, so only vmalloc users gain: modules, BPF, and large
   vmalloc_huge() hash tables. It is still a fairly contained patch.

3. userfaultfd write-protect (HAVE_ARCH_USERFAULTFD_WP)

   Not specific to huge pages, but it covers hugetlbfs too. It needs a
   spare PTE bit plus a swap PTE bit. Bits <31:17> look free, but check.

Not worth it
============

- HVO (HUGETLB_PAGE_OPTIMIZE_VMEMMAP): needs SPARSEMEM_VMEMMAP, which
  Alpha does not have. The saving is about 24KB per 4MB page (0.6%).
- ARCH_WANT_HUGE_PMD_SHARE: only applies to pages the size of a PUD
  entry.
- hugetlb_cma / gigantic pages: 4MB is below MAX_PAGE_ORDER, so the buddy
  allocator already handles it.

Other alpha TODOs in Documentation/features/
============================================

- vm: THP (item 1), huge-vmap (item 2), ioremap_prot, ELF-ASLR, and TLB
  (batched unmap flush).
- Everything else: jump-labels, kprobes/uprobes, stackprotector, KASAN,
  kcov, eBPF JIT, queued spinlocks/rwlocks, context-tracking and NUMA
  balancing.

Plan: send the current series first, then prototype item 1 on megalith.
The mm change decides whether it is viable, so float it on linux-mm
early and point to arm64 contpte as precedent.