There are 2 options to do this in Google colab. Option 1 (Easiest) With this option, we can just install the relevant package and then run the model. First we need to install the required packages using the following command:- ! pip install -U llama-cpp-python Next, we will try to get the model using the following command # !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id= "unsloth/Qwen3-0.6B-GGUF" , filename= "Qwen3-0.6B-IQ4_NL.gguf" , ) We can see the model being downloaded:- And then we will run the following command to test this model llm.create_chat_completion( messages = [ { "role" : "user" , "content" : "What is the capital of France?" } ] ) And the output will look something like this:- Option 2 llama.cpp is a powerful inference engine and if you wanted to get it running in google colab, you co...
Comments