As ARM64 users, our goal was to find an alternative to Ollama capable of leveraging the Hexagon NPU built into Qualcomm Snapdragon X processors on Windows ARM64.
During our evaluation, we discovered the AnythingLLM project, which allowed us to run local models and initially appeared to solve our requirements.
While investigating a specific issue together with the AnythingLLM developers, we discovered that AnythingLLM uses GenieX as its inference backend on Windows ARM64. In other words, GenieX is the component actually responsible for running local models.
For our systems, we only need LLM inference capabilities. We do not require RAG, document management, workspaces, a graphical user interface, or a corporate knowledge base. Our requirements are straightforward: run local models, support conversational interactions, provide streaming responses, and expose a well-known API.
Once we understood the underlying architecture, we realized that AnythingLLM was not the solution we were truly looking for. Instead, it helped us discover GenieX, the component that actually met our needs.
As a result, we decided to remove the intermediate layers and integrate directly with GenieX. This approach reduces potential points of failure, lowers memory consumption, simplifies management and maintenance, decreases architectural complexity, and minimizes dependencies when running local AI inference workloads.
AnythingLLM remains an excellent solution for document-centric and RAG-based use cases. However, if your primary objective is to expose local models through an API for integration into your own applications and services, using GenieX directly may be a simpler and more efficient approach.
Run it from command line as geniex serve or as geniex serve --host 0.0.0.0:18181 if you want to make it available through all network interfaces.
One thought on “Using GenieX Instead of AnythingLLM on Windows ARM64”