empero-ai/Qwen3.8-2B-Distill-GGUF Q4_K_M
This is a good model and really fast too. Using the smaller size Q4_K_M and found out that it this version does not know how to call tool compare to its larger version empero-ai/Qwen3.8-4B-Distill-GGUF Switching to the Q_8 version aint any better. It still does not know how to call tools - the context given is way larger about 32k. Then switching BF16 gives a context of 26.9k which is good but still unable to make any tool calling