Dedicated model · Available as managed deployment

Request a Llama 3.2 deployment on your own DGX Spark

Meta's small Llama 3.2 models — 1B and 3B instruction-tuned — for fast, cheap text tasks that do not need a large model: classification, extraction, summaries, on-device style workloads on a server. Validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only — an OpenAI-compatible endpoint on hardware only you use, operated by AxForge in the EU.

eu-es-1 · Málaga Available as managed deployment Text Quoted per deployment llama-3-2.axforge.ai
Request deploymentTalk to an engineerSign in €0.69/hour on demand · €0.66/hour by the week · €0.62/hour by the month · €0.55/hour by the year, excl. VAT Hardware rental plus a managed service quoted per deployment — both confirmed in writing before anything is billed.

Why AxForge

Why Llama 3.2 as a managed deployment

Small, fast, cheap1B and 3B parameters: thousands of requests per hour on a single dedicated machine, with latency a large model cannot match.
The right size for the jobRouting, tagging, extraction, short summaries, guardrails — the tasks where a small model at high volume beats a big one at any price.
Same endpoint as the big onesOpenAI-compatible on your own hardware, so the small model slots into a pipeline next to Llama 3.1 or Qwen without a different SDK.

Specifications

What you get

ModelLlama 3.2 — meta-llama
ModalitiesText
Sizes1.2B, 3.2B
LicenceOpen, with conditions — llama3.2; AxForge deploys under it and tells you what applies
HardwareNVIDIA DGX Spark (GB10, 128 GB unified memory) — owned and operated by AxForge
Rental termHour, week, month or year
Hardware pricing€0.69/hour on demand · €0.66/hour by the week · €0.62/hour by the month · €0.55/hour by the year, excl. VAT
Managed serviceQuoted per deployment
RegionMálaga, Spain (eu-es-1)

Full details, benchmarks and FAQ on the Llama 3.2 page. Prices exclude VAT.

How it works

From sign-in to running

1Request deployment — describe your traffic, context needs and rental term.
2You receive the configuration, hardware rental and managed-service price in writing before anything is billed.
3AxForge deploys Llama 3.2 on a dedicated DGX Spark reserved for you.
4Point your OpenAI SDK at your own endpoint with the model name you receive.
5Adjust the term — hour, week, month or year — as your workload settles.

Request deployment or sign in to start.

FAQ

Llama 3.2 — common questions

Is Llama 3.2 on the AxForge serverless API?

Not on the serverless API — it is available as a managed deployment: validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. The serverless API serves Qwen3.8 27B.

What is Llama 3.2 good for?

High-volume, low-latency text tasks: classification, extraction, routing, short summaries and guardrails. For general chat quality look at Llama 3.1 or Qwen3.

Can Llama 3.2 run alongside another model?

Yes — a dedicated DGX Spark has room for a small model next to a large one; AxForge configures both behind your endpoint.

How fast is it on your hardware?

AxForge publishes only numbers it measures itself, and has not benchmarked this model on its nodes yet. For quality benchmarks, see the official model card.

What does EU Llama 3.2 hosting cost?

Hardware by the hour, week, month or year; the managed service is quoted per deployment — both confirmed in writing before anything is billed.

Ready for Llama 3.2 on your own machine?

Request deployment Sign in Talk to an engineer

Explore

More from AxForge

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms