Quantized Models
A working catalog of quantization targets covered by TyloQuant and HeadWiseKV across model architectures, precision budgets, and local hardware.
Model weights, licenses, and available releases are defined by the Hugging Face repositories. This page is not a universal performance claim.Different structures require different quantization strategies
These profiles show current adaptation and validation directions rather than forcing every model into one fixed precision template.
DENSE / ROUTED MoE
Qwen3.6
Long context and expert-level precision allocation
Covers the 27B long-context serving path and routed-MoE exploration through mixed-format quantization and head-wise KV-cache allocation.
Qwen3.6-27B is part of the public HeadWiseKV reproduction path. Precision tiers still require validation against target hardware and quality goals.
DENSE / LONG CONTEXT
Qwen3.5
Layer-wise and KV-head sensitivity analysis
Uses the 9B model as a calibration reference for long-context cache and quantization error across window sizes, precisions, and runtime settings.
Public materials include sensitivity analysis and aggregate results. Refer to the repository for model files and complete test conditions.
ARCHITECTURE ADAPTATION
Gemma 4
Structure-aware mixed-format quantization
Assigns encoding formats and precision budgets around internal weight distributions while validating TyloQuant format selection and packed execution.
Adaptation status changes with upstream architectures and runtime support. Unreleased weights are never presented here as downloads.
EXPERT-AWARE QUANTIZATION
DeepSeek V4 Flash
Heterogeneous expert containers and low-bit budgets
Explores per-expert formats and bitrate budgets for MoE models so quantization follows expert sensitivity and the actual inference path.
This is a research adaptation target. Formal versions, licenses, and supported environments are defined by the release page.
Quantization is more than choosing Q4 or Q8
TyloQuant connects encoding formats, precision allocation, and runtime execution so layers, compute groups, and experts can receive different budgets.
A base family for regular low-bit integer encodings.
Vector and grouped encodings for varied weight distributions and quality budgets.
An extended path for weights that need finer-grained precision.
Stores heterogeneous expert formats and bitrates in one MoE tensor.
This page describes research coverage; repositories define availability
Downloads, commercial use, and future API access depend on upstream licenses, validation state, and release repositories. Every deployment still requires testing for target hardware, context length, quality metrics, and concurrency.
浙公网安备33011002019019号