CacheWarden
Keeps assistant prompt caches warm while you step away, so you don't pay to rebuild them.
An open-source coding agent for VS Code, built for people who run their own models.
Forge loads GGUF models through llama.cpp and manages them properly: start, keep loaded, share across editor windows. Cloud models and the Claude Code and Codex CLI agents are there when you choose to add them.


A private, searchable memory of every AI coding session you have had.
HalluScribe reads the session files your AI tools already write, summarises each one with a local model, and keeps a readable Markdown archive on your disk. Nothing is uploaded and there is no API bill.

We are adding first-class support for NVIDIA's open Nemotron models to Forge, with a tuned tool-calling profile for local agentic coding on NVIDIA GPUs. Next, we plan to fine-tune Nemotron for Greek, using the same pipeline and native-speaker data behind our Greek Gemma 4 model. Greek-language benchmarks will be published in the Forge repository.
Keeps assistant prompt caches warm while you step away, so you don't pay to rebuild them.
A coordination board for several AI agents at once: file claims, a shared event feed, and stop and pause for all.
A desktop ring that shows how full your model's context window is, and when quality is likely to drop.
An offline learning companion for children aged 6–11, answering in their own language. No internet, no subscription.
Greek is one of the EU's 24 official languages and one of the least served by open models. We release ours publicly, in formats that run on a single consumer GPU.
Gemma 4 E4B fine-tuned for Greek speech understanding and Greek text, trained on 3,217 native voice recordings across 17 categories and 2,476 curated question–answer pairs. Ships as a single GGUF for Ollama or llama.cpp.
A high-quality open Greek text-to-speech voice for Piper, recorded by a native speaker. Released under CC BY-NC 4.0.
evolv is an independent software studio in Greece. We work on one problem: making open-weight models genuinely useful for real work, on hardware people already own.
Raw inference speed is llama.cpp's job, and it does it well. The hard part is everything above the runtime: tool calls that don't break halfway, context that doesn't quietly overflow, file edits you can undo. That layer is what we build.
Questions, collaboration, or feedback on our tools.