LZ4 compression, xxHash3 integrity, AES-256-GCM encryption — for arbitrary byte payloads.
Features · Quick Start · FFI · Security · Architecture
cachekit-core transforms byte payloads: compress them, verify their integrity, encrypt them. Bytes in, bytes out.
| Component | What it does |
|---|---|
| ByteStorage | &[u8] → LZ4 compress → xxHash3 checksum → Vec<u8> envelope |
| Encryption | &[u8] → AES-256-GCM encrypt → Vec<u8> ciphertext |
| Key Derivation | Master key → HKDF-SHA256 → derived key per tenant/domain |
Tip
For decorator-based caching in Python, see cachekit-py.
| Feature | Description | Default |
|---|---|---|
compression |
LZ4 compression via lz4_flex |
✅ |
checksum |
xxhash-rust integrity verification |
✅ |
encryption |
AES-256-GCM (ring on native, aes-gcm on wasm32) + HKDF-SHA256 (hkdf) |
❌ |
ffi |
C header generation | ❌ |
# Cargo.toml - defaults only
[dependencies]
cachekit-core = "0.1"
# With encryption
[dependencies]
cachekit-core = { version = "0.1", features = ["encryption"] }
# For C FFI development
[dependencies]
cachekit-core = { version = "0.1", features = ["ffi", "encryption"] }use cachekit_core::ByteStorage;
// Create storage with default format
let storage = ByteStorage::new(None);
// Store data (compresses + checksums automatically)
let data = b"Hello, cachekit!";
let envelope = storage.store(data, None)?;
// Retrieve data (decompresses + verifies checksum)
let (retrieved, format) = storage.retrieve(&envelope)?;
assert_eq!(data.as_slice(), retrieved.as_slice());use cachekit_core::{ByteStorage, ZeroKnowledgeEncryptor, derive_domain_key};
use zeroize::Zeroizing; // add the zeroize crate to your Cargo.toml
// Derive tenant-isolated key from master secret
// From your secret manager or CACHEKIT_MASTER_KEY, hex-decoded to 32 raw bytes.
// Never hard-code it, and never pass the hex string's bytes.
// Zeroizing wipes each key from memory when it is dropped.
let master_key = Zeroizing::new(load_master_key_from_secret_manager()?);
let tenant_key = Zeroizing::new(derive_domain_key(
master_key.as_slice(),
"cache", // domain separation
b"tenant-12345", // tenant isolation
)?);
// Encrypt sensitive data
let encryptor = ZeroKnowledgeEncryptor::new()?;
let plaintext = b"sensitive user data";
let aad = b"tenant-12345"; // Additional authenticated data
let ciphertext = encryptor.encrypt_aes_gcm(plaintext, tenant_key.as_slice(), aad)?;
// Decrypt (fails if AAD doesn't match)
let decrypted = encryptor.decrypt_aes_gcm(&ciphertext, tenant_key.as_slice(), aad)?;
assert_eq!(plaintext.as_slice(), decrypted.as_slice());Important
Key Management: Never hardcode keys. Use environment variables or a secrets manager. The CACHEKIT_MASTER_KEY environment variable is the recommended approach.
Full Pipeline: Compress → Encrypt → Store
use cachekit_core::{ByteStorage, ZeroKnowledgeEncryptor, derive_domain_key};
fn cache_sensitive_data(
data: &[u8],
master_key: &[u8],
tenant_id: &str,
) -> Result<Vec<u8>, Box<dyn std::error::Error>> {
// Step 1: Compress + checksum
let storage = ByteStorage::new(None);
let compressed = storage.store(data, None)?;
// Step 2: Derive tenant key
let tenant_key = derive_domain_key(master_key, "cache", tenant_id.as_bytes())?;
// Step 3: Encrypt compressed envelope
let encryptor = ZeroKnowledgeEncryptor::new()?;
let ciphertext = encryptor.encrypt_aes_gcm(
&compressed,
&tenant_key,
tenant_id.as_bytes(),
)?;
Ok(ciphertext)
}Build with FFI feature to generate include/cachekit.h:
cargo build --release --features ffiThis produces:
target/release/libcachekit_core.{so,dylib,dll}— Shared libraryinclude/cachekit.h— C header file
Example C Usage
#include "cachekit.h"
#include <stdio.h>
int main() {
// Create storage handle
CachekitByteStorage* storage = cachekit_byte_storage_new(NULL);
// Store data
const uint8_t data[] = "Hello from C!";
uint8_t* envelope = NULL;
size_t envelope_len = 0;
CachekitError err = cachekit_byte_storage_store(
storage, data, sizeof(data) - 1, NULL, &envelope, &envelope_len
);
if (err != CACHEKIT_OK) {
printf("Store failed: %d\n", err);
return 1;
}
// Retrieve data
uint8_t* retrieved = NULL;
size_t retrieved_len = 0;
err = cachekit_byte_storage_retrieve(
storage, envelope, envelope_len, &retrieved, &retrieved_len
);
// Cleanup
cachekit_byte_storage_free(storage);
cachekit_free_buffer(envelope);
cachekit_free_buffer(retrieved);
return 0;
}Compile:
gcc -o example example.c -L target/release -lcachekit_core -I include┌─────────────────────────────────────────────────────────────────┐
│ Security Architecture │
├─────────────────────────────────────────────────────────────────┤
│ │
│ Master Key ──┬──► HKDF-SHA256 ──► Tenant Key A │
│ │ │
│ ├──► HKDF-SHA256 ──► Tenant Key B │
│ │ │
│ └──► HKDF-SHA256 ──► Tenant Key N │
│ │
│ Each tenant key provides: │
│ • Cryptographic isolation (compromise one ≠ compromise all) │
│ • Domain separation (cache vs auth vs sessions) │
│ • Master-key rotation via decrypt-only keyring (grace window) │
│ │
└─────────────────────────────────────────────────────────────────┘
| Property | Implementation |
|---|---|
| Encryption | AES-256-GCM (AEAD) via ring on native, aes-gcm on wasm32 |
| Key Derivation | HKDF-SHA256 (RFC 5869) via hkdf |
| Integrity | xxhash-rust (xxHash3-64) |
| Nonce Safety | Counter-based + random IV (no reuse) |
| Memory Safety | zeroize on drop for all key material |
| Timing Safety | Constant-time AEAD tag verification: ring on native, aes-gcm on wasm32 |
Warning
Nonce Counter: Each ZeroKnowledgeEncryptor instance supports 2³² encryptions before requiring rotation. The FFI layer returns CACHEKIT_ROTATION_NEEDED at 2³¹ operations as an early warning.
Decompression Bomb Protection
All decompression operations enforce:
| Limit | Value | Purpose |
|---|---|---|
| Max uncompressed size | 512 MB | Memory exhaustion prevention |
| Max compressed size | 512 MB | Input validation |
| Max compression ratio | 1000x | Decompression bomb detection |
Malicious payloads claiming original_size: 500GB with 100 bytes of data are rejected before decompression.
Envelope Decode Bounds
retrieve() and validate() run a header-only structural pre-scan over the
envelope bytes before MessagePack decoding: nesting deeper than 100 levels,
headers declaring more elements or bytes than the input can back, the reserved
marker 0xc1 and truncated input are all rejected before decoding, without
allocating in proportion to any declared length. A rejection is
ByteStorageError::DeserializationFailed with the message prefix
decode pre-scan: . The same walk is public as
check_msgpack_structure(bytes, max_depth) for callers that decode untrusted
MessagePack themselves; it returns a MsgpackStructureError whose Display is
the bare reason, with no prefix. See SECURITY.md.
cachekit-core/
├── src/
│ ├── lib.rs # Public API exports
│ ├── byte_storage.rs # LZ4 + xxHash3 storage envelope
│ ├── msgpack_bounds.rs # Structural pre-scan (public check_msgpack_structure), run before the envelope decode
│ ├── checksum.rs # Standalone xxHash3 checksum/verify primitive (feature = "checksum")
│ │
│ ├── encryption/ # (feature = "encryption")
│ │ ├── mod.rs # Module exports
│ │ ├── core.rs # AES-256-GCM implementation
│ │ ├── key_derivation.rs # HKDF-SHA256 + tenant isolation
│ │ └── keyring.rs # Multi-key decrypt keyring (master-key rotation)
│ │
│ └── ffi/ # (feature = "ffi")
│ ├── mod.rs # FFI exports
│ ├── error.rs # C-compatible error codes
│ ├── handles.rs # Opaque handle management
│ ├── byte_storage.rs # ByteStorage FFI bindings
│ └── encryption.rs # Encryption FFI bindings
│
├── include/
│ └── cachekit.h # Generated C header
│
├── benches/
│ ├── hot_path.rs # Criterion wall-clock suite (make bench)
│ └── perf_ir.rs # Per-op instruction counts and their gate (make perf-ir)
│
├── fuzz/ # Fuzzing targets (16 targets)
│ └── fuzz_targets/
│
└── tests/ # Integration & property tests
Benchmarks on Apple M2 Max (64KB payload, compressible data):
| Operation | Throughput | Notes |
|---|---|---|
| LZ4 compress | ~15 GB/s | Highly compressible data |
| LZ4 decompress | ~37 GB/s | |
| xxHash3-64 | ~36 GB/s | 19x faster than Blake3 |
| AES-256-GCM encrypt | ~6 GB/s | ARM Crypto Extensions |
| AES-256-GCM decrypt | ~6 GB/s | ARM Crypto Extensions |
1KB payload (per-call overhead visible)
| Operation | Throughput |
|---|---|
| LZ4 compress | ~2 GB/s |
| LZ4 decompress | ~14 GB/s |
| xxHash3-64 | ~10 GB/s |
| AES-256-GCM encrypt | ~3.6 GB/s |
| AES-256-GCM decrypt | ~4.4 GB/s |
Tip
Hardware acceleration is auto-detected. ARM64 uses ARM Crypto Extensions; x86-64 uses AES-NI.
The figures above come from examples/bench_throughput.rs on highly compressible data. The Criterion suite in benches/hot_path.rs needs the encryption feature (make bench, or cargo bench --features encryption; a plain cargo bench skips it). It runs the ByteStorage roundtrip on three corpora side by side: byte_storage/roundtrip (synthetic ramp, kept for history), byte_storage/roundtrip_msgpack (realistic msgpack records, about 0.38 LZ4 ratio at 64 KB) and byte_storage/roundtrip_incompressible. The bench profile keeps symbols, so callgrind and perf attribute cost to functions.
make perf-ir counts the instructions each hot path executes per call and fails when a case costs 1% or more above its committed budget. It warns from 0.2%. Wall clock on a shared machine moves by far more than 1% between identical runs, and Criterion picks its iteration count from wall time; instruction counts at a fixed iteration count are exact. It needs valgrind (apt install valgrind).
| Case | What one call does |
|---|---|
store/<size>, retrieve/<size> |
ByteStorage::store / retrieve on the realistic msgpack payload hot_path uses |
prescan/<size> |
check_msgpack_structure on the envelope, the pre-scan retrieve runs first |
encrypt/<size>, decrypt/<size> |
ZeroKnowledgeEncryptor AES-256-GCM |
keyring_decrypt/<size> |
Keyring::decrypt_indexed, which re-derives the tenant key (HKDF) on every call |
tenant_keyring_decrypt/<size> |
TenantKeyring::decrypt_indexed, key derived once at for_tenant |
hkdf |
derive_domain_key (takes no payload, so it has no size) |
Sizes are 64 B, 1 KiB and 64 KiB. One case at a time: make perf-ir ARGS="--case store/1024".
Method (benches/perf_ir.rs): each case runs in its own process under valgrind --tool=cachegrind --cache-sim=no, once with 1,000 calls and once with none, after the same setup and one warm-up call. Ir/op is (Ir[1000] - Ir[0]) / 1000, so process start, setup and exit cancel. The measured process gets an empty environment, so neither the shell nor a VALGRIND_OPTS changes what is counted.
Reproducibility: two runs of the same build agree exactly on every case. Between builds, the heap moves with the size of the binary, and alignment-dependent code in AES-GCM and memcpy follows it: encrypt/1024 moves by up to 0.19%, every other case by at most 0.02%. The counts also depend on the compiler, the locked dependencies, the CPU features the binary is compiled for and the ones valgrind passes through. Budgets are keyed by architecture, OS and any compiled-in CPU features (a -C target-cpu=x86-64-v3 build gets its own set), and each set records the rustc and valgrind versions it was taken with; the gate prints a note when yours differ. Other RUSTFLAGS and profile overrides are not detected, so measure the default build.
Budgets live in benches/perf_ir_baselines.json. After a change that makes a path cheaper, make perf-ir-update ratchets its budget down. It never raises one: an increase you mean takes make perf-ir-update ARGS=--allow-increase. Ratcheting works only on the rustc and valgrind the budgets were recorded with; on a new toolchain, re-record every case with make perf-ir-update ARGS=--allow-increase (no --case). Instruction counts ignore cache misses and branch mispredictions, so a claimed wall-clock win still needs a wall-clock A/B.
# Run all tests
cargo test --all-features
# Run with specific feature
cargo test --features encryption
# Property-based tests
cargo test --all-features -- --include-ignored proptest
# Fuzzing (requires cargo-fuzz)
cd fuzz && cargo fuzz run byte_storage_corrupted_envelopeSee fuzz/README.md for comprehensive fuzzing documentation.
tests/wire_format_vectors.rs byte-verifies this crate against the canonical
ByteStorage envelope vectors from
cachekit-io/protocol
(test-vectors/wire-format.json). Since protocol 1.1 the fixture pins two
compressed_data encodings — legacy array-of-integers vectors and their
*_bin twins (msgpack bin, canonical for 1.1+ writers): every vector in
both sets must decode to the exact payload bytes, while re-encode
byte-identity is asserted against the set matching this crate's current
writer encoding — msgpack bin, since the protocol 1.1 serde_bytes
writer flip (checksum deliberately stays array-of-ints per the
protocol's normative scope exclusion). Legacy envelopes remain readable
forever.
tests/dual_decode.rs proves both reader shapes accept both encodings,
including bin16/bin32 width headers. The fixture is vendored at
tests/vectors/wire-format.json and integrity-pinned by sha256 — to update
it, re-copy from the protocol repo and change the pinned hash in the same
commit.
Since protocol 1.3 the fixture also carries six reject_vectors, which
tests/wire_format_vectors.rs drives through retrieve(), asserting the
error the protocol names for each. src/read_allocation_probe.rs (unit tests
only) bounds what the size-cap and ratio reads allocate, with a positive
control. tests/wire_format_constructed.rs runs the fixture's constructed
32-bit ratio-product vector natively and, in CI, on wasm32-unknown-unknown
under wasm-bindgen-test.
tests/decode_bounds_vectors.rs drives every reject and accept vector in the
protocol's test-vectors/decode-bounds.json (vendored the same way, at
tests/vectors/decode-bounds.json) through retrieve(). Each reject vector
must fail with the pre-scan's decode pre-scan: message prefix; failing
somewhere inside the decoder does not count.
This crate requires Rust 1.85 or later (Edition 2024).
User-facing docs in this repository follow CacheKit's shared rule on what belongs in them:
What belongs in these docs.
prek install (or pre-commit install) sets up hooks that reject internal references in README
files, docs/ and commit messages.
MIT License — see LICENSE for details.