Using Decision Models to Make Small LLMs More Efficient

During the development of our Mr. Crab AI agent framework, we realized that many interactions inside an agentic workflow are actually decision-making tasks with only a limited number of possible outcomes.

Small language models are not necessarily bad at reasoning. More often, they are being asked to perform the wrong task. When we request structured output, they may struggle to consistently produce well-formatted XML, JSON, or YAML responses.

When running AI agents on mini PCs, laptops, ARM-based systems, or machines without powerful GPUs, every generated token matters. Token generation consumes both computation time and resources, which directly impacts responsiveness and efficiency.

Many agentic frameworks repeatedly ask their language model questions such as:

  • Should a tool be executed?
  • Should the workflow continue?
  • Is the generated answer valid?
  • Is more information required?
  • Which module should be executed next?
  • Has the conversation reached its conclusion?

A traditional language model may respond with something like:

{
  "continue": true,
  "reason": "The user is asking for additional information"
}

or even:

{
  "continue": true
}

when the framework only needs the answer:

true

To improve reliability when working with tiny and small language models, we progressively reduced the complexity of the prompts sent to the model.

Instead of asking the model to generate structured XML, JSON, or YAML responses, many internal workflow steps were transformed into simple classification, routing, and yes/no decision tasks.

This naturally led us to explore decision models such as Jev AI and tev1, which are specifically designed for this type of workload. Decision models focus on selecting predefined outcomes rather than generating free-form text, making them particularly well suited for workflow orchestration and routing tasks.

As a result, we are gradually moving from the traditional workflow:

User
  │
  ▼
LLM
 ├─ Search?
 ├─ Call Tool?
 ├─ Retry?
 ├─ Continue?
 └─ Generate Response?

towards a more specialized architecture:

User
  │
  ▼
Decision Model
 ├─ Search?
 ├─ Call Tool?
 ├─ Continue?
 ├─ Retry?
 └─ Escalate?
        │
        ▼
 Small or Large LLM
        │
        ▼
 Generated Response

This approach allows the decision model to handle workflow orchestration while reserving the language model for what it does best: generating natural language responses.

It is important to note that decision models are not replacements for traditional language models. They are complementary components. In environments where decision models are not available, such as some local inference backends, these tasks can still be performed by a standard LLM. However, when a decision model is available, it can significantly reduce the amount of text generation required and improve overall efficiency.

For organizations running local AI workloads, reducing resource consumption is not only about lowering costs. It is also about making advanced AI workflows accessible on ordinary hardware, including mini PCs, laptops, ARM devices, and other resource-constrained systems.

Decision models represent an interesting evolution in agent architecture. Instead of asking a language model to make every workflow decision and generate every response, we can separate decision-making from language generation.

By using the right model for the right task, small LLMs become more reliable, more efficient, and more practical for local AI deployments.

Why We Integrated Both Ollama and AnythingLLM for Local AI in Community Management

At Communities of Neighbors Management System, privacy and transparency are at the heart of everything we build. That’s why we’ve integrated support for local large language models (LLMs), giving community administrators the ability to generate professional announcements and notifications without sending sensitive data to external services.

We started with Ollama, a powerful local LLM runner that makes it easy to deploy models directly on personal computers. Ollama ensures that announcements—such as water pipe bursts, lighting outages, or maintenance notices—can be drafted quickly and securely, with all data staying inside the community’s environment.

But we didn’t stop there. We also integrated AnythingLLM, which brings unique advantages for users on Windows 11 ARM64 devices powered by Qualcomm processors. Unlike Ollama, AnythingLLM supports NPUs (Neural Processing Units) natively, unlocking hardware acceleration and improved performance on modern ARM64 systems. This means faster inference, lower energy consumption, and smoother experiences for administrators working on Qualcomm-powered PCs.

Additionally, AnythingLLM offers RAG (Retrieval-Augmented Generation) capabilities. Administrators can connect shared documents within AnythingLLM, allowing the LLM to reference community-specific files when generating announcements. This makes notifications more accurate, contextual, and tailored to the needs of each building or neighborhood.

By supporting both Ollama and AnythingLLM, we give communities the freedom to choose the local AI solution that best fits their hardware and workflow. Whether it’s CPU-based inference with Ollama or NPU-accelerated generation with AnythingLLM, administrators can rely on our platform to deliver professional communication while safeguarding resident privacy.

Local AI is not just about performance—it’s about trust. With Ollama and AnythingLLM, Communities of Neighbors Management System empowers administrators to manage communication responsibly, securely, and efficiently.

Join our social network channels for quick updates: