llama.cr
Crystal bindings for llama.cpp, a C/C++ implementation of LLaMA, Falcon, GPT-2, and other large language models.
The version in shard.yml corresponds to the compatible llama.cpp build number.
For example, shard version 0.10809.1 targets llama.cpp build b10809, which
is the stable release v0.4.0.
This project is under active development and may change rapidly.
Features
- Low-level bindings to the llama.cpp C API
- High-level Crystal wrapper classes for easy usage
- Memory management for C resources
- Simple text generation interface
- Backend capability checks and configurable GPU offloading
- Advanced sampling methods (Min-P, Typical, Mirostat, etc.)
- Batch processing for efficient token handling
- KV cache management for optimized inference
- State saving and loading
Installation
Install llama.cpp first, then add this shard.
1. Install llama.cpp
macOS (Homebrew)
brew install llama.cpp
export LLAMA_LIB_DIR="$(brew --prefix llama.cpp)/lib"
Linux (prebuilt release matching this shard version)
VERSION="$(shards version)"
BUILD="$(echo "$VERSION" | sed -E 's/^0\.([0-9]+)\.[0-9]+$/\1/')"
LLAMA_BUILD="b${BUILD}"
curl -L "https://github.com/ggml-org/llama.cpp/releases/download/${LLAMA_BUILD}/llama-${LLAMA_BUILD}-bin-ubuntu-x64.tar.gz" -o llama.tar.gz
tar -xzf llama.tar.gz
sudo cp llama-${LLAMA_BUILD}/*.so* /usr/local/lib/
sudo ldconfig
2. Add to your project
dependencies:
llama:
github: kojix2/llama.cr
version: 0.<build>.<patch>
Then run:
shards install
Pin an exact version because llama.cpp updates can include breaking changes between build numbers.
3. Build and run
Linux:
export LLAMA_LIB_DIR=/path/to/llama.cpp/lib
LIBRARY_PATH="$LLAMA_LIB_DIR" crystal build examples/simple.cr \
--link-flags "-L$LLAMA_LIB_DIR -Wl,-rpath,$LLAMA_LIB_DIR -lllama -lggml"
LD_LIBRARY_PATH="$LLAMA_LIB_DIR" ./simple --model models/tiny_model.gguf
macOS:
export LLAMA_LIB_DIR=/path/to/llama.cpp/lib
LIBRARY_PATH="$LLAMA_LIB_DIR" crystal build examples/simple.cr \
--link-flags "-L$LLAMA_LIB_DIR -Wl,-rpath,$LLAMA_LIB_DIR -lllama -lggml"
DYLD_LIBRARY_PATH="$LLAMA_LIB_DIR" ./simple --model models/tiny_model.gguf
If backend auto-detection fails in newer llama.cpp builds, set GGML_BACKEND_PATH to a backend shared library file (not a directory), for example:
export GGML_BACKEND_PATH="$LLAMA_LIB_DIR/libggml-cpu-haswell.so"
Advanced setup
Build from source:
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
VERSION="$(shards version ..)"
BUILD="$(echo "$VERSION" | sed -E 's/^0\.([0-9]+)\.[0-9]+$/\1/')"
LLAMA_BUILD="b${BUILD}"
git checkout "${LLAMA_BUILD}"
mkdir build && cd build
cmake .. && cmake --build . --config Release
sudo cmake --install . && sudo ldconfig
Example for local development/tests:
MODEL_PATH=/path/to/model.gguf \
LIBRARY_PATH="$LLAMA_LIB_DIR" \
LD_LIBRARY_PATH="$LLAMA_LIB_DIR" \
GGML_BACKEND_PATH="$LLAMA_LIB_DIR/libggml-cpu-haswell.so" \
crystal spec
Obtaining GGUF Model Files
You'll need a model file in GGUF format. For testing, smaller quantized models (1-3B parameters) with Q4_K_M quantization are recommended.
Popular options:
Usage
Basic Text Generation
require "llama"
response = Llama.generate(
"/path/to/model.gguf",
"Once upon a time",
max_tokens: 100,
temperature: 0.8
)
puts response
For a typed result with a finish reason, token usage, stop sequences, or
streaming, use the additive Llama.complete API:
options = Llama::GenerationOptions.new(max_tokens: 100, stop: ["\n\n"])
result = Llama.complete(
"/path/to/model.gguf",
"Once upon a time",
options
)
puts result.text
puts result.finish_reason
puts result.usage.tokens_per_second
Llama.generate continues to return a String. Llama.complete returns a
Generation; both use the same checked generation path.
Typed model and context construction policy can be supplied without managing the resources manually:
result = Llama.complete(
"/path/to/model.gguf",
"Once upon a time",
options,
model_options: Llama::ModelOptions.new(gpu_layers: -1),
context_options: Llama::ContextOptions.new(context_size: 4096_u32)
)
Streaming and Stateful Sessions
Use a Session to reuse one context. Chunks contain valid UTF-8 and stop text is
buffered so it is not emitted unless include_stop is enabled.
Llama::Model.open("/path/to/model.gguf") do |model|
model.session do |session|
result = session.generate("Once upon a time") do |chunk|
print chunk.text
STDOUT.flush
end
puts "\n#{result.finish_reason}"
end
end
Session keeps a canonical transcript between calls. Call reset to start a
new sequence. snapshot, restore, save, and load validate the model and
native version before changing that transcript. Snapshots store canonical text,
not native KV-cache bytes. transcript_token_count counts logical transcript
tokens; the compatibility method used_tokens is also a tokenizer result and
does not report native KV occupancy. Only one generation may use a session at a
time.
A streamed GenerationChunk is a text-delivery unit. UTF-8 buffering and stop
prefix detection mean it need not correspond one-to-one with a sampled token.
Its token is the most recent token that triggered delivery and its index is
a monotonic emission index, including final flushes. Usage timing is
end-to-end wall time and therefore includes work performed by the streaming
block. Cooperative cancellation is checked between prompt batches and generated
tokens, so latency is bounded by one native call rather than being instantaneous.
Backend Capabilities and GPU Offloading
GPU offloading depends on the linked llama.cpp build and the backends available at runtime. Check capabilities before selecting accelerator-specific settings:
puts Llama.gpu_offload_supported?
puts Llama.mmap_supported?
puts Llama.mlock_supported?
puts Llama.rpc_supported?
With GPU offloading available, the convenience API can offload the model and context operations:
raise "GPU offloading is unavailable" unless Llama.gpu_offload_supported?
response = Llama.generate(
"/path/to/model.gguf",
"Once upon a time",
n_gpu_layers: -1,
offload_kqv: true,
op_offload: true
)
n_gpu_layers: -1 requests all model layers; 0 keeps them on the CPU.
offload_kqv controls KQV operations and the KV cache, while op_offload
controls host tensor operations. These options do not add GPU support to a
CPU-only llama.cpp build. The defaults are n_gpu_layers: 0,
offload_kqv: false, and op_offload: false.
The same context settings are available when managing resources directly:
Llama::Model.open("/path/to/model.gguf", n_gpu_layers: -1) do |model|
model.context(offload_kqv: true, op_offload: true) do |context|
puts context.generate("Once upon a time")
end
end
Lazy Model Loading
llama.cpp can load eligible model tensors on demand. The default is
Llama::LazyMode::AUTO, which lazily loads marked tensors larger than 4 GiB.
Llama::Model.open("/path/to/model.gguf", lazy_mode: Llama::LazyMode::ON) do |model|
model.context do |context|
puts context.generate("Once upon a time")
end
end
Use Llama::LazyMode::OFF to always read complete tensors up front.
The convenience API accepts the same setting:
response = Llama.generate(
"/path/to/model.gguf",
"Once upon a time",
lazy_mode: Llama::LazyMode::ON
)
Resource Lifetime
Native-backed objects provide an idempotent free method. Prefer the block APIs
shown above for deterministic cleanup. For longer-lived resources, call free
in an ensure block and release dependencies before their owners:
model = Llama::Model.new("/path/to/model.gguf")
context = model.context
begin
puts context.generate("Once upon a time")
ensure
context.free
model.free
end
Release samplers and adapters before contexts, and contexts before models.
Samplers added to a SamplerChain are released with the chain.
Backend Lifetime
Llama.init is called automatically when a model or context is created, so most
applications do not need to call it manually.
Llama.uninit is optional and usually not needed. It is intended only for
controlled teardown after all Llama::Model and Llama::Context instances have
been finalized. Calling it while models or contexts are still alive raises an
error, because their finalizers may still need the llama.cpp backend.
Saved State Compatibility
llama.cpp b10809 updates the session and sequence-state file formats. Session or state files written by b10566 are not guaranteed to load with this version; recreate them after upgrading.
Advanced Sampling
require "llama"
Llama::Model.open("/path/to/model.gguf") do |model|
model.context do |context|
Llama::SamplerChain.open do |chain|
chain.add(Llama::Sampler::TopK.new(40))
chain.add(Llama::Sampler::MinP.new(0.05, 1))
chain.add(Llama::Sampler::Temp.new(0.8))
chain.add(Llama::Sampler::Dist.new(42))
result = context.generate_with_sampler("Write a short poem about AI:", chain, 150)
puts result
end
end
end
Chat Conversations
require "llama"
Llama::Model.open("/path/to/model.gguf") do |model|
model.chat(system: "You are a helpful assistant.") do |chat|
result = chat.ask("Hello, who are you?") do |chunk|
print chunk.text
end
puts "\n#{result.finish_reason}"
end
end
Chat history is committed transactionally. A cancelled turn is not committed
unless commit_partial: true is requested. b10809 recognizes predefined chat
template shapes; it is not a general Jinja evaluator. Pass a recognized template
explicitly when the model does not provide one. Chat#clear removes every
message, including the initial system message. A second turn, clear, or close
during an active turn raises BusyError.
Embeddings
require "llama"
Llama::Model.open("/path/to/model.gguf") do |model|
model.embedder(pooling: Llama::Pooling::Mean) do |embedder|
vector = embedder.embed("Hello, world!", normalize: true)
vectors = embedder.embed_all(["one", "two"], normalize: true)
puts "Embedding dimension: #{vector.size}"
end
end
Embedder owns a dedicated embedding context, copies native vectors before the
next call, and preserves input order in native multi-sequence batches.
Utilities
System Info
info = Llama.runtime_info
puts "llama.cr #{info.wrapper_version} expects #{info.expected_build}"
puts "loaded llama.cpp #{info.reported_version} with #{info.backend_count} backends"
puts info.system_info
Tokenization Utility
Llama::Model.open("/path/to/model.gguf", vocab_only: true) do |model|
puts Llama.tokenize_and_format(model.vocab, "Hello, world!", ids_only: true)
end
Examples
The examples directory contains sample code demonstrating various features:
simple.cr- Basic text generationminimal.cr- Minimal use of the high-level generation APIchat.cr- Chat conversations with modelsstreaming.cr- UTF-8-safe streamed generation with a sessionembedding.cr- Single-batch normalized sentence embeddingstokenize.cr- Tokenization and vocabulary featuresserver.cr- HTTP streaming server (usesexamples/shard.yml)
API Documentation
See kojix2.github.io/llama.cr for full API docs.
API Layers
- Convenience API:
Llama.generateandContext#generateretain their existing string-returning behavior. - Typed helpers:
GenerationOptions,Session,Chat,Embedder, andSampling::Planadd structured results and managed workflows. - Advanced API:
Context,Batch,Memory,State, and manual samplers expose native concepts. Borrowed views are valid only while their owner remains open; copy pointer-backed data before another native call. Callers composing these low-level operations are responsible for synchronization; the high-levelBusyErroroperation guard does not make arbitrary raw call sequences atomic. In particular, logits and embedding pointers may be invalidated by the nextdecode/encode, and every borrowed model/context view becomes invalid when its owning context or model is closed. - Raw API:
require "llama/raw"exposesLlama::LibLlama. Its structs, symbols, and pointer lifetimes track the pinned upstream build and may change between shard releases.
Llama.check_compatibility! can explicitly check the version reported by the
library. It is not enforced during initialization, allowing other llama.cpp
versions to be tested at the user's own risk. Stable packages of the supported
release report 0.4.0, while the official b10809 release archives report
0.4.0-dev. The C API does not report the exact build number, so package pins
and ABI checks are still needed when strict compatibility is required.
Custom Llama.log_set callbacks are experimental. On the pinned b10809 build,
model loading and multithreaded decode callbacks were observed on the calling
thread. Callback exceptions are contained at the C boundary and can be retrieved
with Llama.take_log_callback_error.
Core Classes
- Llama::Model - Represents a loaded LLaMA model
- Llama::Context - Handles inference state for a model
- Llama::Vocab - Provides access to the model's vocabulary
Llama::Session- Reusable typed and streaming generationLlama::Chat- Transactional conversation historyLlama::Embedder- Safe single and batched sentence embeddings- Llama::Batch - Manages batches of tokens for efficient processing
- Llama::Memory - Controls KV cache memory and related operations
- Llama::State - Handles saving and loading model state
- Llama::SamplerChain - Combines multiple sampling methods
Samplers
- Llama::Sampler::TopK - Keeps only the top K most likely tokens
- Llama::Sampler::TopP - Nucleus sampling (keeps tokens until cumulative probability exceeds P)
- Llama::Sampler::Temp - Applies temperature to logits
- Llama::Sampler::Dist - Samples from the final probability distribution
- Llama::Sampler::MinP - Keeps tokens with probability >= P * max_probability
- Llama::Sampler::Typical - Selects tokens based on their "typicality" (entropy)
- Llama::Sampler::Mirostat - Dynamically adjusts sampling to maintain target entropy
- Llama::Sampler::Penalties - Applies penalties to reduce repetition
Development
See DEVELOPMENT.md for development guidelines.
This software is primarily created through AI-generated code.
Do you need commit rights?
- If you need commit rights to my repository or want to get admin rights and take over the project, please feel free to contact @kojix2.
- Many OSS projects become abandoned because only the founder has commit rights to the original repository.
Contributing
- Fork it (https://github.com/kojix2/llama.cr/fork)
- Create your feature branch (
git checkout -b my-new-feature) - Commit your changes (
git commit -am 'Add some feature') - Push to the branch (
git push origin my-new-feature) - Create a new Pull Request
License
This project is available under the MIT License. See the LICENSE file for more info.