← BACK TO PORTFOLIO

CASE STUDY / AI INFRASTRUCTURE

SYJ-LLM

A C++17 local AI runtime built around llama.cpp and GGUF, with a local model registry, memory-aware inference, CLI workflows, localhost API foundation and a fine-tuning pipeline.

ACTIVE · PHASE 5C++17 · llama.cpp · GGUF · CMake · ARM64

Why it matters

The flagship systems project: the portfolio demonstrates ownership of the runtime boundary rather than only consuming a hosted model API.

Architecture

SYSTEM / VERIFIED ARCHITECTURE LENS
RuntimeNative C++17 inference core
Model layerGGUF registry + memory-aware admission
InterfaceCLI workflows + localhost API foundation
BuildCMake + pinned llama.cpp
TargetOffline / constrained ARM64 environments
Public repository This page intentionally avoids invented performance metrics, customer counts or production claims.

Evidence / implementation boundary

01Current evidence

C++17 runtime work, pinned llama.cpp integration, GGUF model registry, memory-aware inference, CLI workflows, localhost API foundation and a Phase 5 fine-tuning pipeline are documented in the repository.

02Validation context

The project is developed with constrained environments in mind, including ARM64/Termux validation supplied by the project owner. This page does not turn that into a universal hardware benchmark.

03Experimental

Fine-tuning and broader runtime capabilities remain active areas of development.

04Limitations

Model performance and memory requirements vary by model, quantization and hardware. Production deployment is not claimed merely from the existence of the runtime.

Roadmap

Continue runtime hardening, reproducible benchmarks, broader model workflows and clearer deployment guidance.

Engineering focus

Problem framing → architecture → implementation → validation → documentation → the next useful version. The portfolio is designed to show decisions and evidence, not just screenshots.