Why HiCache Matters
SGLang HiCache extends the traditional RadixAttention with a three-tier hierarchical KV caching system that dramatically improves performance for long-context and multi-turn conversation scenarios. By intelligently managing KV caches across GPU memory, host memory, and external storage backends, HiCache addresses the fundamental capacity bottleneck that limits cache hit rates in conventional systems.L1 and L2 are private to a single inference instance; only L3 can be shared. Host memory cannot be pooled across instances or across hosts, not even for two instances on the same node. Raising
--hicache-ratio or --hicache-size only enlarges the instance’s own private L2. Cross-instance reuse is the job of L3, which needs --hicache-storage-backend and a backend configured for the scope you want: the file backend defaults to the node-local /tmp/hicache, whereas mooncake, hf3fs, nixl and aibrix reach cluster scope when every instance shares the same namespace. See Tier Sharing Scope.Configuration Guidelines
Core HiCache Parameters
Command
- Besides configuring
--hicache-storage-backendat startup, SGLang also supports runtime attach/detach of the HiCache storage backend (no restart required) via HTTP admin endpoints. See Runtime Attach/Detach HiCache Storage Backend.
Key Configurations with Storage Backends Enabled
Memory Layout Optimization
Command
page_first: Only compatible withkernelI/O backend, automatically switches tolayer_firstwithdirectbackendpage_first_direct: Specifically designed fordirectI/O backend with optimized memory organization
Heterogeneous TP Support (GQA/MHA models)
HiCache storage supports cross-cluster KV reuse when different deployments use different TP sizes (for example,tp=4 and tp=8) and share the same storage backend namespace.
Use tp_lcm_size in --hicache-storage-backend-extra-config:
Command
- Set
tp_lcm_sizeto the least common multiple (LCM) of all TP sizes that will share the same HiCache storage. - For MHA models with Mooncake and
page_headlayout, HiCache will split head shards based ontp_lcm_sizeto make keys reusable across heterogeneous TP deployments. - If all clusters use the same TP size, this option is not needed.
Prefetch Policies
Command
Integration with PD Disaggregation
HiCache works seamlessly with PD Disaggregation. You can choose between two configurations:- Prefill-only HiCache: Enable HiCache only on Prefill nodes, allowing KV cache sharing among Prefill instances
- Full HiCache with async offloading: Enable HiCache on Prefill nodes and async KV cache offloading on Decode nodes, allowing Prefill nodes to reuse KV caches from Decode nodes in multi-turn dialogue scenarios
Command
Deployment with Local Filesystem Backends
file is the minimal reference backend: unbounded, staged, serial file I/O.
For a local NVMe or filesystem tier use fast_file. It keeps the same on-disk
page format and adds direct host-buffer I/O for page-first layouts, parallel
reads, atomic page publication, and background LRU eviction:
Command
storage_dir in a namespace derived from the model name,
parallel layout, and host pool layout, so one root can be shared by several
deployments. Without storage_dir, fast_file uses the file backend’s
SGLANG_HICACHE_FILE_BACKEND_STORAGE_DIR and then /tmp/hicache; the
namespace directory keeps its pages apart from file pages in the same root.
max_size bounds one namespace as tracked by one process (MLA ranks share a
namespace and rank 0 owns eviction); min_free_space is a floor for the whole
filesystem. Both are optional. Eviction state and the optional
enable_metadata_cache are process-local, and a completed write means close
plus rename rather than fsync. The module docstring of
sglang/srt/mem_cache/storage/fast_file/fast_file_store.py lists every
extra-config key.
Deployment with HF3FS
Here is an example of deploying DeepSeek-R1 with HiCache-HF3FS. For more details, see the HF3FS Documentation.Command
Deployment with Mooncake
Here is an example of deploying Qwen3-235B-A22B-Instruct-2507 with Mooncake. For more details, see the Mooncake Documentation.Command
Custom Storage Backend Integration
To integrate a new storage backend:-
Implement three core methods:
get(key): Retrieve value by keyexists(key): Check key existenceset(key, value): Store key-value pair
- Register your backend: Add your storage backend to the HiCache BackendFactory
Dynamic Backend Loading
Alternatively, you can use dynamic loading to avoid hard-coding your backend in the repository:Command
--hicache-storage-backend: Set todynamic--hicache-storage-backend-extra-config: JSON configuration with:backend_name: Custom backend identifiermodule_path: Python module path to your implementationclass_name: Your HiCache implementation class nameinterface_v1: 0 (disable) or 1 (enable) to control usage of batch_get_v1 and batch_set_v1 methods
Community and Support
- GitHub Issues: Report bugs and feature requests
- Slack Channel: Join community discussions in #sgl-kv-cache-store
- Documentation: Refer to storage backend-specific guides
This document will be continuously updated based on community feedback and new features. Contributions and suggestions are welcome!
