MODEL CATALOG / QUANTIZED INFERENCETYLOGI / PUBLIC INDEX

Quantized Models

A working catalog of quantization targets covered by TyloQuant and HeadWiseKV across model architectures, precision budgets, and local hardware.

Model weights, licenses, and available releases are defined by the Hugging Face repositories. This page is not a universal performance claim.
4Model families
16TyloQuant encodings
0.84–8.30BPW research range
MODEL PROFILES / 01

Different structures require different quantization strategies

These profiles show current adaptation and validation directions rather than forcing every model into one fixed precision template.

01In validation

DENSE / ROUTED MoE

Qwen3.6

Long context and expert-level precision allocation

Covers the 27B long-context serving path and routed-MoE exploration through mixed-format quantization and head-wise KV-cache allocation.

Q4_K_MQ6_KQ8_0NINTM v2

Qwen3.6-27B is part of the public HeadWiseKV reproduction path. Precision tiers still require validation against target hardware and quality goals.

02Calibration reference

DENSE / LONG CONTEXT

Qwen3.5

Layer-wise and KV-head sensitivity analysis

Uses the 9B model as a calibration reference for long-context cache and quantization error across window sizes, precisions, and runtime settings.

Q8_0PPLHeadWiseKVGGUF

Public materials include sensitivity analysis and aggregate results. Refer to the repository for model files and complete test conditions.

03Adapting

ARCHITECTURE ADAPTATION

Gemma 4

Structure-aware mixed-format quantization

Assigns encoding formats and precision budgets around internal weight distributions while validating TyloQuant format selection and packed execution.

NINTNVQNPQNEPQ

Adaptation status changes with upstream architectures and runtime support. Unreleased weights are never presented here as downloads.

04Adapting

EXPERT-AWARE QUANTIZATION

DeepSeek V4 Flash

Heterogeneous expert containers and low-bit budgets

Explores per-expert formats and bitrate budgets for MoE models so quantization follows expert sensitivity and the actual inference path.

NINTM v2Mixed BPWCUDAC++ Runtime

This is a research adaptation target. Formal versions, licenses, and supported environments are defined by the release page.

FORMAT SYSTEM / 02

Quantization is more than choosing Q4 or Q8

TyloQuant connects encoding formats, precision allocation, and runtime execution so layers, compute groups, and experts can receive different budgets.

NINT

A base family for regular low-bit integer encodings.

NVQ / NPQ

Vector and grouped encodings for varied weight distributions and quality budgets.

NEPQ

An extended path for weights that need finer-grained precision.

NINTM v2

Stores heterogeneous expert formats and bitrates in one MoE tensor.

RELEASE BOUNDARY / 03

This page describes research coverage; repositories define availability

Downloads, commercial use, and future API access depend on upstream licenses, validation state, and release repositories. Every deployment still requires testing for target hardware, context length, quality metrics, and concurrency.

[ INTERACT TO CREATE LIFE ]