Using Decision Models to Make Small LLMs More Efficient

During the development of our Mr. Crab AI agent framework, we realized that many interactions inside an agentic workflow are actually decision-making tasks with only a limited number of possible outcomes.

Small language models are not necessarily bad at reasoning. More often, they are being asked to perform the wrong task. When we request structured output, they may struggle to consistently produce well-formatted XML, JSON, or YAML responses.

When running AI agents on mini PCs, laptops, ARM-based systems, or machines without powerful GPUs, every generated token matters. Token generation consumes both computation time and resources, which directly impacts responsiveness and efficiency.

Many agentic frameworks repeatedly ask their language model questions such as:

  • Should a tool be executed?
  • Should the workflow continue?
  • Is the generated answer valid?
  • Is more information required?
  • Which module should be executed next?
  • Has the conversation reached its conclusion?

A traditional language model may respond with something like:

{
  "continue": true,
  "reason": "The user is asking for additional information"
}

or even:

{
  "continue": true
}

when the framework only needs the answer:

true

To improve reliability when working with tiny and small language models, we progressively reduced the complexity of the prompts sent to the model.

Instead of asking the model to generate structured XML, JSON, or YAML responses, many internal workflow steps were transformed into simple classification, routing, and yes/no decision tasks.

This naturally led us to explore decision models such as Jev AI and tev1, which are specifically designed for this type of workload. Decision models focus on selecting predefined outcomes rather than generating free-form text, making them particularly well suited for workflow orchestration and routing tasks.

As a result, we are gradually moving from the traditional workflow:

User
  │
  ▼
LLM
 ├─ Search?
 ├─ Call Tool?
 ├─ Retry?
 ├─ Continue?
 └─ Generate Response?

towards a more specialized architecture:

User
  │
  ▼
Decision Model
 ├─ Search?
 ├─ Call Tool?
 ├─ Continue?
 ├─ Retry?
 └─ Escalate?
        │
        ▼
 Small or Large LLM
        │
        ▼
 Generated Response

This approach allows the decision model to handle workflow orchestration while reserving the language model for what it does best: generating natural language responses.

It is important to note that decision models are not replacements for traditional language models. They are complementary components. In environments where decision models are not available, such as some local inference backends, these tasks can still be performed by a standard LLM. However, when a decision model is available, it can significantly reduce the amount of text generation required and improve overall efficiency.

For organizations running local AI workloads, reducing resource consumption is not only about lowering costs. It is also about making advanced AI workflows accessible on ordinary hardware, including mini PCs, laptops, ARM devices, and other resource-constrained systems.

Decision models represent an interesting evolution in agent architecture. Instead of asking a language model to make every workflow decision and generate every response, we can separate decision-making from language generation.

By using the right model for the right task, small LLMs become more reliable, more efficient, and more practical for local AI deployments.

Community Fees & Membership Dues

We’ve just published our last version of the Communities of Neighbors Management System adding the “Community Fees & Membership Dues” section.

Community administrators and managers will be able to add fees and dues.

Community Fees & Membership Dues

Create the fees and dues to collect using criteria by tags and apartment dimensions.

Create Fee Set

Your neighbors will receive a notification per email.

Community Fees & Membership Dues emails

With a convenient link for your neighbors to review their pending fees and dues.

Pending Community Fees & Membership Dues

And let them pay online or offline depending on what payment gateways have been configured for your community by the community administrators.

Pay Community Fees & Membership Dues

MrCrab: Building a Lightweight Agentic AI Framework for Small Models

Introduction Over the past year, agentic AI frameworks have grown rapidly, but many of them are designed with large models in mind. This creates friction when developers want to run smaller models locally, either for cost efficiency, privacy, or speed. After experimenting with OpenClaw, NanoClaw, PicoClaw, and Nanobot, we decided to build our own agent: MrCrab.

Why Small Models Matter

  • Local deployment without dependency on cloud quotas.
  • Lower hardware requirements, making AI accessible to more teams.
  • Faster iteration cycles for debugging and prototyping.
  • Privacy and compliance advantages when data never leaves your infrastructure.

Design Principles of MrCrab

  • Tool Registry: Instead of injecting a massive tool list into every prompt, MrCrab allows the agent to query tools dynamically by keyword.
  • Hybrid Memory: Recent turns are kept in context, older ones are summarized, and full logs are stored in persistent memory for retrieval on demand.
  • Backend Flexibility: MrCrab integrates with Ollama, AnythingLLM, and any provider compatible with the OpenAI API.
  • Lightweight Prompts: Optimized for small models like Gemma4:e2b, Qwen3.5:2B, Granite4:1B.

Implementation Highlights

  • Modular architecture written with simplicity in mind.
  • Debugging and logging designed to be transparent.
  • Easy integration with local or cloud‑based LLMs.

Lessons Learned

  • Large prompts and tool lists overwhelm small models.
  • Timeout handling must be explicit when working with local inference.
  • Summarization should be progressive, not premature.

Future Work

  • Extending MrCrab to real business use cases, such as community management systems.
  • Adding support for multi‑agent collaboration.
  • Exploring long‑context training for small models.

MrCrab is our attempt to make agentic AI practical for small models. By focusing on lightweight prompts, dynamic tool discovery, and hybrid memory, we believe it can bridge the gap between experimental frameworks and production‑ready agents.

MrCrab is written in PHP, without third‑party libraries. This design choice minimizes supply chain risks and ensures that the agent can be deployed in a secure and portable way. Developers can run MrCrab locally with minimal setup, while still benefiting from integrations with Ollama, AnythingLLM, and OpenAI‑compatible APIs.

Why We Integrated Both Ollama and AnythingLLM for Local AI in Community Management

At Communities of Neighbors Management System, privacy and transparency are at the heart of everything we build. That’s why we’ve integrated support for local large language models (LLMs), giving community administrators the ability to generate professional announcements and notifications without sending sensitive data to external services.

We started with Ollama, a powerful local LLM runner that makes it easy to deploy models directly on personal computers. Ollama ensures that announcements—such as water pipe bursts, lighting outages, or maintenance notices—can be drafted quickly and securely, with all data staying inside the community’s environment.

But we didn’t stop there. We also integrated AnythingLLM, which brings unique advantages for users on Windows 11 ARM64 devices powered by Qualcomm processors. Unlike Ollama, AnythingLLM supports NPUs (Neural Processing Units) natively, unlocking hardware acceleration and improved performance on modern ARM64 systems. This means faster inference, lower energy consumption, and smoother experiences for administrators working on Qualcomm-powered PCs.

Additionally, AnythingLLM offers RAG (Retrieval-Augmented Generation) capabilities. Administrators can connect shared documents within AnythingLLM, allowing the LLM to reference community-specific files when generating announcements. This makes notifications more accurate, contextual, and tailored to the needs of each building or neighborhood.

By supporting both Ollama and AnythingLLM, we give communities the freedom to choose the local AI solution that best fits their hardware and workflow. Whether it’s CPU-based inference with Ollama or NPU-accelerated generation with AnythingLLM, administrators can rely on our platform to deliver professional communication while safeguarding resident privacy.

Local AI is not just about performance—it’s about trust. With Ollama and AnythingLLM, Communities of Neighbors Management System empowers administrators to manage communication responsibly, securely, and efficiently.

Join our social network channels for quick updates:

Affiliate Program at the Communities of Neighbors

Our Affiliate Program is now live! 🎉
You can create your account, get your affiliate ID, and start earning commissions by inviting communities to use our management system.

👉 Check out your affiliate panel and start today!

#AffiliateProgram #CommunityManagement #SmartCommunity #EarnCommissions #NeighborhoodManagement

Communities of Neighbors Management System

Discounts at Communities of Neighbors

Looking for discount coupons for your new community of neighbors created in our Communities of Neighbors Management System? Say no more!

🏘️ Build a stronger neighborhood with Communities of Neighbors! 🏘️

Manage announcements, track issues, chat with residents, & more – all in one secure platform. Perfect for HOAs, property managers, & engaged communities.

✨ Special Offer! ✨ Get 6% off your first month with code: SOCIAL-6989B03F (Valid for 1 week).

➡️ Learn more & sign up: https://communitiesofneighbors.lucentinian.com/

#Community #HOA #Neighborhood #PropertyManagement #ResidentEngagement #LocalCommunity #CommunityBuilding

Don’t miss our channel in the Fediverse where we’re posting coupon codes weekly!

https://social.lucentinian.com/profile/communitiesofneighbors

Are you an X / Twitter user? Then find them in our channel at

Lucentinian Works Co Ltd (@ehehdadaltd) / X

Tasks added to the issues

Every community of neighbors face issues that you can track with our Communities of Neighbors Management System.

We’ve added task tracking for a fine grain management.

Set up a short description as the task title, set the completion status, create subtasks, set the ideal and the actual dates and set dependencies.

Generate combined Gantt charts with the progress of the tasks of the issue.

Or separated in ideal dates and actual dates for better clarification, especially with large number of tasks.

Track de progress deviation:

Create announcements with the progress of the work to solve the issues of your community of neighbors and get the content for your announcement assisted by your favorite AI provider for a more professional result.

Assign your announcement to an electronic board to community the progress of the issue to all the neighbors in your community.

Get your announcements published in your electronic boards across the community of neighbors that you’re managing:

🏡 Introducing Our New Neighborhood Management System

We’re proud to announce the launch of our latest service: the Communities of Neighbors Management System — a digital platform designed to simplify communication, coordination, and participation in residential communities.

Whether you’re a property manager, building administrator, or resident, this system helps you stay connected and informed through:

  • 📢 Announcements
  • 📝 Surveys and forms
  • 💬 Secure real-time chat
  • 🧾 Organized dashboards for community management

Our goal is to make neighborhood life more transparent, collaborative, and efficient.

🔗 Learn More

You can explore the full details of the system here:
👉 https://communitiesofneighbors.lucentinian.com/

🎁 Follow Us for Weekly Coupons

To celebrate the launch, we’ll be publishing weekly discount coupons and updates through our official channels:

Follow us to stay informed and enjoy exclusive offers!

💬 Let’s Build Better Communities Together

We believe that strong communities start with better tools. If you’re managing a residential building or looking to improve communication with your neighbors, this platform is for you.

We’re excited to grow this project with your feedback and participation.

More of 1 year of AI generated jokes

We’re still producing new jokes almost every day. When it happened is because the flow has been broken due OS updates or because the AI entere in a loop trying to validate that the news which is used to generate the jokes are suitable for a joke. Don’t hesitate in visiting us at https://comics.lucentinian.com !

Merry Christmas!!

🎄🎅🏻 Ho ho hoooo! 🎅🏻🎄

Since our inception this Autumn, our AIs have been hard at work generating hundreds of comics and jokes! Now, we need your help to choose the best ones! 🌟

Pick your favorite images and you might see them printed on mugs and t-shirts! ☕👕

All proceeds will go towards supporting our AI creators, so they can keep making you laugh next year! 😂💫

Visit us!

https://comics.lucentinian.com

Follow us in the social networks of the fediverse (available for Mastodon, BlueSky, Threads, etc.):

https://social.lucentinian.com/profile/comics

Add us to your RSS feeds:

https://comics.lucentinian.com/rss.xml

Support us at:

Do you want to help us by donating some cryptocurrency?