Free LLM ModelNemotron 3 Ultra 550B
MODEL CALLING GUIDE · NVIDIA NIM

Use Nemotron 3 Ultra 550B for free via NVIDIA NIM

NVIDIA · 1,000,000 token context on this channel · Coding, Reasoning, Tool calling. This guide documents a specific API channel and has not run an inference request.

Docs checked 2026-09-29Inference not testedOpenAI Chat Completions
01 / FREE TERMS

Check the free conditions first

Free typeFree quota
Price in this tierFree endpoint for eligible NVIDIA API keys; paid partner deployment is separate
QuotaUp to 40 requests/minuteNVIDIA NIM free-tier provider snapshot; account limits can differ
RegistrationRequired
Payment cardNot required for this free tier
Phone verificationRequiredThe current freeLLM provider snapshot marks phone verification as required.
Region / eligibilityNot confirmed for every accountCheck the platform signup and model page for your region.
Free periodNo announced end date · subject to changeNo announced end date; free endpoint availability and limits can change.
02 / GET STARTED

Five steps to a first response

  1. 1

    Create or sign in to the NVIDIA Developer account

    Open the NVIDIA API Catalog and sign in. The public quickstart describes this as a free developer account; the catalog may ask for phone verification.

    Open NVIDIA API Catalog ↗
  2. 2

    Create an API key

    Open Settings → API Keys, generate a key, and keep it in your local shell. This site never collects it.

    Open NVIDIA API Keys ↗
  3. 3

    Use this exact NIM model

    Copy the NVIDIA Base URL and exact Model ID below. The free endpoint is separate from paid partner deployment options.

  4. 4

    Send one small request

    Set NVIDIA_API_KEY locally and run the cURL example. It uses NVIDIA's OpenAI-compatible Chat Completions endpoint.

  5. 5

    Watch the free limit

    The current directory snapshot records up to 40 RPM from the freeLLM provider page. Recheck the model page or dashboard before relying on that number; NVIDIA can change account eligibility and limits.

    Read the official quickstart ↗
03 / CONNECTION

Copy these exact public values

SDK Base URLhttps://integrate.api.nvidia.com/v1
Full HTTP endpointhttps://integrate.api.nvidia.com/v1/chat/completions
Model IDnvidia/nemotron-3-ultra-550b-a55b
Protocol / authenticationOpenAI Chat Completions · Bearer $NVIDIA_API_KEY
Extra required fieldsNone documented for this text request
04 / FIRST CALL

Minimal cURL request

Run this on your own computer. Set NVIDIA_API_KEY in your local shell first. The page never collects or sends your key. The sample requests at most 256 output tokens and contains no paid fallback.

Terminal
# Set NVIDIA_API_KEY in your own terminal first. Do not paste it into this site.
curl 'https://integrate.api.nvidia.com/v1/chat/completions' \
  -H "Authorization: Bearer $NVIDIA_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"model":"nvidia/nemotron-3-ultra-550b-a55b","messages":[{"role":"user","content":"Reply with one short greeting."}],"max_tokens":256}'

A 401/403 can indicate a key or account issue; 404 can indicate a changed model ID; 429 usually points to a limit. Check NVIDIA NIM's own error details before changing plans.

05 / SOURCES

Where these details came from

Documentation checked 2026-09-29. Inference request: not run. Prices, quotas and access can change.

← Back to Free LLM Model