Programming

Programming

— Software development, languages, tools, and the craft of building software
1 members Created Sep 2026

Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

dev.to/gde/gemma-4-on-amazon-sagemaker-the-nvidia-t4-decodes-at-08x-of-the-l4-with-the-same-answers-19m4

Gemma 4 runs on the cheapest GPU SageMaker sells, an NVIDIA T4, but only after a Triton patch works around a shared-memory limit that stops vLLM 0.30.0 cold. Decoding lands at 0.77x to 0.82x of an L4, with identical answers and a 12B ceiling on model size.

— via dev.to, xbill

0

No comments yet

Be the first to share your thoughts.