Posts

commander using llm model openai/XHToken/Spark-X2.5-4B-GGUF

Image
When running my commander app with XHToken/Spark-X2.5-4B-GGUF the result is pretty good. The model is able to follow instruction and extracts out the right container image to run and the commands too.   This is my codebase for setting up my model XHToken/Spark-X2.5-4B-GGUF. import os from dotenv import load_dotenv from google . adk . agents import LlmAgent from google . adk . models . lite_llm import LiteLlm from google . adk . tools import google_search # Create a LiteLLM model pointing to your local server model = LiteLlm (     model = "openai/XHToken/Spark-X2.5-4B-GGUF" ,     api_base = "http://localhost:8888/v1" ,   # Your local server     api_key = "sk-unsloth-4d0a1b198bd177a2a72ee1954585342a" ,     temperature = 0.0 ,     extra_body = { "chat_template_kwargs" : { "enable_thinking" : False }}, ) from app . prompt import ROOT_AGENT_INSTRUCTION from app . tools . container_toolset import (   ...

commander - using ornith-ai/Ornith-1.0-9B-GGUF as the model and documenting the results

Image
 After trying out a couple of model, I decides to see if using a different model would sway the results differently. I am using ornith-ai/Ornith-1.0-9B-GGUF. model here.  Noticed that I have added "openai" in front of the model name otherwise it throw an error message. import os from dotenv import load_dotenv from google . adk . agents import LlmAgent from google . adk . models . lite_llm import LiteLlm from google . adk . tools import google_search # Create a LiteLLM model pointing to your local server model = LiteLlm (     model = "openai/ornith-ai/Ornith-1.0-9B-GGUF" ,     api_base = "http://localhost:8888/v1" ,   # Your local server     api_key = "sk-unsloth-4d0a1b198bd177a2a72ee1954585342a" ,     temperature = 0.0 ) from app . prompt import ROOT_AGENT_INSTRUCTION from app . tools . container_toolset import (     container_image_finder_tool , run_container_command_tool ) load_dotenv () root_agent ...

Quickstart with laya

Image
To get started with laya.  Laya is a free, open-source AI model designed for fast, focused text classification. It makes simple decisions—such as routing emails, prioritising tickets, or identifying refund requests—by choosing from predefined options and providing a confidence score. In short: Laya is built for fast, low-cost AI decisions rather than generating text. To get started, we have to install laya and creating your python environment. Next we going to use it to decide if the ticket belongs to which department. This is an approach where it uses the model with your local computing resources. There's another option for you use Laya studio which requires an API but you have to sign up first. import laya from laya import Router # Initialize router with preloading (avoids swap delay) router = Router ( preload = True ) # Define complex state ticket = {     "ticket_id" : "TCK-8821" ,     "customer" : "enterprise_user" ,     "su...

keycloak setting up mcp cimd profile

Image
In this example, we will be setting up our mcp authentication using Keycloak CIMD using vscode desktop.  Let's get our keycloak instance up and running.  docker run -p 8080:8080 -e KC_FEATURES=cimd -e KC_BOOTSTRAP_ADMIN_USERNAME=admin -e KC_BOOTSTRAP_ADMIN_PASSWORD=admin quay.io/keycloak/keycloak:latest start-dev Then setup follow these steps here to configure vscode desktop Setting up the client profile for VS Code desktop Navigate to Realm Settings → Client Policies → Profiles tab. Click Create client profile . Give the profile a name such as vscode-cimd-profile and click Save . Click Add executor and select client-id-metadata-document from the list. Configure the executor with the following options: Allow http scheme : OFF Trusted domains : vscode.dev , 127.0.0.1 , code.visualstudio.com (This option is applied not only to the client_id URL but also to the URL-valued properties of the Client ID Metadata Document, such as client_uri , logo_uri , tos_uri , policy_...

Google ADK creating custom tool

  Creating a custom tool with Google ADK is pretty much the same defining a normal python function.  Here is an example code:- As you can see from the code here get_weather_report() is a custom tool that we can define and wrap it around with FunctionTool.   import asyncio from google . adk . agents import Agent from google . adk . tools import FunctionTool from google . adk . runners import Runner from google . adk . sessions import InMemorySessionService from google . genai import types APP_NAME = "weather_sentiment_agent" USER_ID = "user1234" SESSION_ID = "1234" MODEL_ID = "gemini-2.0-flash" # Tool 1 def get_weather_report ( city : str ) -> dict :     """Retrieves the current weather report for a specified city.     Returns:         dict: A dictionary containing the weather information with a 'status' key ('success' or 'error') and a 'report' key with the weather details if succ...

AggregatedValue is a reserved word in Application Insights

Apparently claude model like to use aggregatedValue as a variable name when creating an alert in Application Insights. After reviewing the code, found out that it is not a good idea  at all.  https://learn.microsoft.com/en-us/azure/azure-monitor/alerts/alerts-create-log-alert-rule#configure-alert-rule-conditions

google adk with local llm agent

Image
 We all wanted our agent to be running 24 hours a day to boost productivity and that normally requires a local LLM. agent. Today we are going to look into setting Qwen3.8 local model using unsloth studio.  First we get unsloth (under the hood it is using llama.cpp) and generate an API token. Then we will proceed to create python env and setup our packages.  uv venv uv init uv add google-adk litellm litert-llm ollama openai Lets' go and create your app name by running the command below :- adk create app  This will create a directory and some basic scaffolding files   Next we will update agent.py with the following code :- from google . adk . agents . llm_agent import LlmAgent from google . adk . models . lite_llm import LiteLlm # Create a LiteLLM model pointing to your local server model = LiteLlm (     model = "openai/empero-ai/Qwen3.8-2B-Distill-GGUF" ,     api_base = "http://localhost:8888/v1" ,   # Your local server   ...

AWS Bedrock agentcore vs harness coding perspective

In bedrock agentcore, we have harness and runtime. We going to look at the code differences trying to invoke them. Whenever possible, ry to use 'bedrock-agentcore' as it is the newer package  atleast for now. Older library typically uses bedrock-agent-runtime . https://docs.aws.amazon.com/boto3/latest/reference/services/bedrock-agent-runtime.html Harness Notice that everything is the same except we need to update harnessArn and we invoke the method invoke_harness(). Notice we are instantiating 'bedrick-agentcore'. There's also a client called 'bedrock-agent-runtime. And when you see the word harness - you know you're in a good seat. :) import boto3 import json import uuid client = boto3.client( 'bedrock-agentcore' , region_name = 'ap-southeast-2' ) session_id = str (uuid.uuid4()) response = client.invoke_harness(     harnessArn = 'your-harness-arn' ,     runtimeSessionId = session_id,     messages = [         {     ...

aws bedrock runtime that integrates with your knowledge based

Image
 First of all we need to setup our knowledge base as shown here where we can specify the source of our knowledge.  Once we created it, we can referenced it in our code, next.  And these are the code that we can use to harness our knowledge base.  import boto3 client = boto3 . client ( "bedrock-agent-runtime" , region_name = "ap-southeast-2" ) # AgenticRetrieveStream - streaming, agent-driven multi-step retrieval response = client .agentic_retrieve_stream(     messages = [{ "role" : "user" , "content" : { "text" : "your query here" }}],     retrievers = [{ "configuration" : { "knowledgeBase" : { "knowledgeBaseId" : "YOUR-KNOWLEDGE-BASE-ID" }}}],     agenticRetrieveConfiguration = { "foundationModelType" : "MANAGED" , "maxAgentIteration" : 5 },     generateResponse = True , ) for event in response [ "stream" ]:     if "respons...