litert-community/gemma-4-26B-A4B-it-litert-lm

🤗 Hugging Face 来源apache-2.0激活 4B32 GBother✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo litert-community/gemma-4-26B-A4B-it-litert-lm ./model-folder
需要做种者 →

litert-community/gemma-4-26B-A4B-it-litert-lm

Main Model Card: google/gemma-4-26B-A4B-it

This model card provides the Gemma 4 26B (A4B) mixture-of-experts model in LiteRT-LM format that is ready for deployment on web and desktop. Please check back here regularly for updates on wider platform support and further functionality improvements. The current LiteRT-LM version supports text; audio, image, and multitoken prediction support will be available in a future update.

Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. This particular Gemma 4 model is a medium size so it is ideal for desktop use cases. By running this model on device, users can have private access to Generative AI technology without even requiring an internet connection.

These models are provided in the .litertlm format for use with the LiteRT-LM framework. LiteRT-LM is a specialized orchestration layer built directly on top of LiteRT, Google’s high-performance multi-platform runtime trusted by millions of Android and edge developers. LiteRT provides the foundational hardware acceleration via XNNPack for CPU and ML Drift for GPU. LiteRT-LM adds the specialized GenAI libraries and APIs, such as KV-cache management, prompt templating, and function calling. This integrated stack is the same technology powering the Google AI Edge Gallery showcase app.

The model is provided as a QAT quantized model with blockwise int4 weights and float activations, specially optimized for running on GPU.

Requirement: Running this model on Desktop requires LiteRT-LM v0.17.1 or later.

Try Gemma 4 26B (A4B)

Build with Gemma 4 26B (A4B) and LiteRT-LM

Ready to integrate this into your product? Get started with LiteRT-LM documentation.

Gemma 4 26B (A4B) Performance on LiteRT-LM

Desktop benchmarks were taken using 8192 prefill tokens and 512 decode tokens with a context length of 16384 tokens via LiteRT-LM CLI. Model size is the size of the file on disk.

Linux

Device                                      Backend Prefill (tokens/sec) Decode (tokens/sec) Init. time (sec) Model size (MB)
NVidia 4090 24GB GPU 4487 86 27.89 15787

macOS

Device                                      Backend Prefill (tokens/sec) Decode (tokens/sec) Init. time (sec) Model size (MB) GPU Memory (MB)
MacBook Pro 2024 (M4 Max) 48GB GPU 1430 77 8.83 15787 ~16190

Windows

Device                                      Backend Prefill (tokens/sec) Decode (tokens/sec) Init. time (sec) Model size (MB)
NVidia 5080 16GB GPU 1676 76 35.40 15787

Web

Web benchmarks were taken in Chrome using 1024 prefill tokens and 256 decode tokens with a context length of 1280 tokens. The model can support up to 195k context length.

Device                                      Backend Quantization Prefill (tokens/sec) Decode (tokens/sec) Time-to-first-token (sec) Model size (MB) GPU Memory (MB) Peak CPU Memory (MB)
MacBook Pro 2024 (M4 Max) 48GB GPU Q4_0 968 51 14.0 15787 ~15000 ~3600