You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

ACE-3-26B-A4B-Preview-260910

APMIC-logo-橫-黑 NVIDIA-NeMo

Model Description

ACE-3-26B-A4B-Preview-260910 is a preview release of APMIC's ACE-3 model family, built for Traditional Chinese (Taiwan) enterprise scenarios and agentic workflows.

The model is based on google/gemma-4-26B-A4B-it, a Mixture-of-Experts model with 26B total parameters and about 4B active parameters per token (the "A4B" in the name). APMIC has further optimized it to strengthen:

  • Traditional Chinese output in Taiwan usage (terminology, phrasing and orthography)
  • Both thinking and non-thinking response modes

Preview notice: this is a preview checkpoint intended for evaluation and feedback. It has not been through a full production release process. Please evaluate it on your own workloads before deploying.


Model Details

  • Developed by: APMIC
  • Model type: Gemma4ForConditionalGeneration (Transformers), sparse MoE (26B total / ~4B active)
  • Base model: google/gemma-4-26B-A4B-it
  • Language(s) (NLP): Traditional Chinese & English
  • Weights precision: bfloat16 (not quantized)
  • License: gemma (Google usage license; gated on Hugging Face)

Usage

messages = [{"role": "user", "content": "請用繁體中文簡單說明什麼是混合專家模型(MoE)。"}]

prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=False,   # True to let the model produce a reasoning block first
)

Transformers

import torch
from transformers import AutoProcessor, AutoModelForImageTextToText

model_id = "APMIC/ACE-3-26B-A4B-Preview-260910"  # replace with the actual repo id

processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
)

messages = [{"role": "user", "content": "請用繁體中文簡單說明什麼是混合專家模型(MoE)。"}]
inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, enable_thinking=False,
    tokenize=True, return_dict=True, return_tensors="pt",
).to(model.device)

out = model.generate(**inputs, max_new_tokens=512)
print(processor.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

Recommended sampling

The bundled generation_config.json uses temperature=1.0, top_p=0.95, top_k=64. These follow the base model's recommended settings.

Serving

The model can be served with any inference engine that supports Gemma 4 (for example vLLM). Because the weights are bfloat16, plan for roughly 52 GB of GPU memory for the weights alone, plus KV cache.


Intended Use and Limitations

Intended for

  • Traditional Chinese (Taiwan) assistants and enterprise applications
  • Function-calling and agent frameworks that need a locally deployable model
  • Research and evaluation of the ACE-3 model family

Limitations

  • This is a preview release; behavior may change in later ACE-3 versions.
  • Like all LLMs, the model can produce incorrect or fabricated content. For high-risk use (financial, legal, medical), keep a human review or enterprise control layer in place.
  • Tool-call outputs should be validated (schema and argument checks) before being executed.
  • Performance on languages other than Traditional Chinese and English has not been specifically optimized.

License

This model is a derivative of Gemma 4 and is distributed under the Gemma Terms of Use. By using it you agree to Google's Gemma Terms of Use and Prohibited Use Policy.

Downloads last month
2
Safetensors
Model size
26B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for APMIC/ACE-3-26B-A4B-Preview

Finetuned
(190)
this model