Gemma 4 QAT
Google公式に発表されたGemmaシリーズの量子化版
性能を維持したまま少メモリ化するために,変数の量子化を学習中に折り込む工夫(量子化対応学習,
QAT
:
Quantization-Aware Training
)が採用されている.
llama.cpp
,
ollama
,
LM Studio
で使用できる.
関連:
Quantization-Aware Training(QAT)
https://gyazo.com/9fff91d9ddda40c4ffbdf6c1c9895475
from:
Gemma 4 with quantization-aware training