Instructions to use nvidia/Llama3-ChatQA-1.5-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nvidia/Llama3-ChatQA-1.5-8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="nvidia/Llama3-ChatQA-1.5-8B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("nvidia/Llama3-ChatQA-1.5-8B") model = AutoModelForCausalLM.from_pretrained("nvidia/Llama3-ChatQA-1.5-8B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use nvidia/Llama3-ChatQA-1.5-8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nvidia/Llama3-ChatQA-1.5-8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/Llama3-ChatQA-1.5-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nvidia/Llama3-ChatQA-1.5-8B
- SGLang
How to use nvidia/Llama3-ChatQA-1.5-8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nvidia/Llama3-ChatQA-1.5-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/Llama3-ChatQA-1.5-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nvidia/Llama3-ChatQA-1.5-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/Llama3-ChatQA-1.5-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use nvidia/Llama3-ChatQA-1.5-8B with Docker Model Runner:
docker model run hf.co/nvidia/Llama3-ChatQA-1.5-8B
π© Report: Legal issue
Nvidia,
Your recent release of the ChatQA-1.5-8B language model raises serious concerns about compliance with the licensing terms for LLaMA 3, the open source model released by Meta AI that your work appears to be derived from.
The LLaMA 3 terms of service explicitly state that any derivative models or products must be named "in a way that clearly indicates their lineage from LLaMA 3." However, the name you have chosen for your 8 billion parameter model, "ChatQA-1.5-8B," makes no mention of its origins from LLaMA 3.
This omission seems to be a direct violation of the licensing requirements Meta has put in place. By obfuscating the connection to LLaMA 3 in the naming, Nvidia appears to be attempting to sidestep the open source licensing terms while benefiting from Meta's publicly released work.
This is a concerning precedent that undermines the principles of open and ethical use of open source AI models. Nvidia needs to be held accountable for this potential breach of license.
I call on Nvidia to immediately either:
- Rename your ChatQA model to properly indicate its lineage from LLaMA 3 in compliance with the terms.
- Or provide a clear, public justification for how your approach does not violate the LLaMA 3 license.
Open and transparent use of open source work is crucial for building trust in the rapidly evolving AI field. Nvidia must take responsibility and respect the licensing terms that enable open sourcing of foundational models like LLaMA 3.
I did it for the Vine.
By Llama 3 license:
"(B) prominently display βBuilt with Meta Llama 3β on a related website, user interface, blogpost, about page, or product documentation. If you use the Llama Materials to create, train, fine tune, or otherwise improve an AI model, which is distributed or made available, you shall also include βLlama 3β at the beginning of any such AI model name."
So 1. Display "Built with Meta Llama 3β on page. 2. Rename model to be prefixed with "Llama 3" as stated in the Llama 3 license.
@zihanliu
Personally I don't think it's realistic for the clause about having the Llama 3 at the beginning of model name's, for Nvidia or anyone else.
People should non-comply and Meta should remove that.
We should welcome open sourcing of all models, in this case we have the weights and the data, that should be commended.
Freegheist
We apologize sincerely for this issue. We just updated the model card and added Llama-3 in the model name.
Can u pls make H100 cheaper? :))))
I think llama3-chatQA-1.5 is also a more marketable name to be honest. The 'llama3' branding gives extra layer of credibility.
Llama3 license is one of the most hostile as it puts such an limitations on how you name your model. I do not understand how you could support it.