llama.cr

test examples docs Lines of Code Static Badge

Crystal bindings for llama.cpp, a C/C++ implementation of LLaMA, Falcon, GPT-2, and other large language models.

The version in shard.yml corresponds to the compatible llama.cpp build number. For example, shard version 0.10809.1 targets llama.cpp build b10809, which is the stable release v0.4.0.

This project is under active development and may change rapidly.

Features

Installation

Install llama.cpp first, then add this shard.

1. Install llama.cpp

macOS (Homebrew)

brew install llama.cpp
export LLAMA_LIB_DIR="$(brew --prefix llama.cpp)/lib"

Linux (prebuilt release matching this shard version)

VERSION="$(shards version)"
BUILD="$(echo "$VERSION" | sed -E 's/^0\.([0-9]+)\.[0-9]+$/\1/')"
LLAMA_BUILD="b${BUILD}"
curl -L "https://github.com/ggml-org/llama.cpp/releases/download/${LLAMA_BUILD}/llama-${LLAMA_BUILD}-bin-ubuntu-x64.tar.gz" -o llama.tar.gz
tar -xzf llama.tar.gz
sudo cp llama-${LLAMA_BUILD}/*.so* /usr/local/lib/
sudo ldconfig

2. Add to your project

dependencies:
  llama:
    github: kojix2/llama.cr
    version: 0.<build>.<patch>

Then run:

shards install

Pin an exact version because llama.cpp updates can include breaking changes between build numbers.

3. Build and run

Linux:

export LLAMA_LIB_DIR=/path/to/llama.cpp/lib
LIBRARY_PATH="$LLAMA_LIB_DIR" crystal build examples/simple.cr \
  --link-flags "-L$LLAMA_LIB_DIR -Wl,-rpath,$LLAMA_LIB_DIR -lllama -lggml"
LD_LIBRARY_PATH="$LLAMA_LIB_DIR" ./simple --model models/tiny_model.gguf

macOS:

export LLAMA_LIB_DIR=/path/to/llama.cpp/lib
LIBRARY_PATH="$LLAMA_LIB_DIR" crystal build examples/simple.cr \
  --link-flags "-L$LLAMA_LIB_DIR -Wl,-rpath,$LLAMA_LIB_DIR -lllama -lggml"
DYLD_LIBRARY_PATH="$LLAMA_LIB_DIR" ./simple --model models/tiny_model.gguf

If backend auto-detection fails in newer llama.cpp builds, set GGML_BACKEND_PATH to a backend shared library file (not a directory), for example:

export GGML_BACKEND_PATH="$LLAMA_LIB_DIR/libggml-cpu-haswell.so"
Advanced setup

Build from source:

git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
VERSION="$(shards version ..)"
BUILD="$(echo "$VERSION" | sed -E 's/^0\.([0-9]+)\.[0-9]+$/\1/')"
LLAMA_BUILD="b${BUILD}"
git checkout "${LLAMA_BUILD}"
mkdir build && cd build
cmake .. && cmake --build . --config Release
sudo cmake --install . && sudo ldconfig

Example for local development/tests:

MODEL_PATH=/path/to/model.gguf \
LIBRARY_PATH="$LLAMA_LIB_DIR" \
LD_LIBRARY_PATH="$LLAMA_LIB_DIR" \
GGML_BACKEND_PATH="$LLAMA_LIB_DIR/libggml-cpu-haswell.so" \
crystal spec

Obtaining GGUF Model Files

You'll need a model file in GGUF format. For testing, smaller quantized models (1-3B parameters) with Q4_K_M quantization are recommended.

Popular options:

Usage

Basic Text Generation

require "llama"

response = Llama.generate(
  "/path/to/model.gguf",
  "Once upon a time",
  max_tokens: 100,
  temperature: 0.8
)
puts response

For a typed result with a finish reason, token usage, stop sequences, or streaming, use the additive Llama.complete API:

options = Llama::GenerationOptions.new(max_tokens: 100, stop: ["\n\n"])

result = Llama.complete(
  "/path/to/model.gguf",
  "Once upon a time",
  options
)
puts result.text
puts result.finish_reason
puts result.usage.tokens_per_second

Llama.generate continues to return a String. Llama.complete returns a Generation; both use the same checked generation path.

Typed model and context construction policy can be supplied without managing the resources manually:

result = Llama.complete(
  "/path/to/model.gguf",
  "Once upon a time",
  options,
  model_options: Llama::ModelOptions.new(gpu_layers: -1),
  context_options: Llama::ContextOptions.new(context_size: 4096_u32)
)

Streaming and Stateful Sessions

Use a Session to reuse one context. Chunks contain valid UTF-8 and stop text is buffered so it is not emitted unless include_stop is enabled.

Llama::Model.open("/path/to/model.gguf") do |model|
  model.session do |session|
    result = session.generate("Once upon a time") do |chunk|
      print chunk.text
      STDOUT.flush
    end
    puts "\n#{result.finish_reason}"
  end
end

Session keeps a canonical transcript between calls. Call reset to start a new sequence. snapshot, restore, save, and load validate the model and native version before changing that transcript. Snapshots store canonical text, not native KV-cache bytes. transcript_token_count counts logical transcript tokens; the compatibility method used_tokens is also a tokenizer result and does not report native KV occupancy. Only one generation may use a session at a time.

A streamed GenerationChunk is a text-delivery unit. UTF-8 buffering and stop prefix detection mean it need not correspond one-to-one with a sampled token. Its token is the most recent token that triggered delivery and its index is a monotonic emission index, including final flushes. Usage timing is end-to-end wall time and therefore includes work performed by the streaming block. Cooperative cancellation is checked between prompt batches and generated tokens, so latency is bounded by one native call rather than being instantaneous.

Backend Capabilities and GPU Offloading

GPU offloading depends on the linked llama.cpp build and the backends available at runtime. Check capabilities before selecting accelerator-specific settings:

puts Llama.gpu_offload_supported?
puts Llama.mmap_supported?
puts Llama.mlock_supported?
puts Llama.rpc_supported?

With GPU offloading available, the convenience API can offload the model and context operations:

raise "GPU offloading is unavailable" unless Llama.gpu_offload_supported?

response = Llama.generate(
  "/path/to/model.gguf",
  "Once upon a time",
  n_gpu_layers: -1,
  offload_kqv: true,
  op_offload: true
)

n_gpu_layers: -1 requests all model layers; 0 keeps them on the CPU. offload_kqv controls KQV operations and the KV cache, while op_offload controls host tensor operations. These options do not add GPU support to a CPU-only llama.cpp build. The defaults are n_gpu_layers: 0, offload_kqv: false, and op_offload: false.

The same context settings are available when managing resources directly:

Llama::Model.open("/path/to/model.gguf", n_gpu_layers: -1) do |model|
  model.context(offload_kqv: true, op_offload: true) do |context|
    puts context.generate("Once upon a time")
  end
end

Lazy Model Loading

llama.cpp can load eligible model tensors on demand. The default is Llama::LazyMode::AUTO, which lazily loads marked tensors larger than 4 GiB.

Llama::Model.open("/path/to/model.gguf", lazy_mode: Llama::LazyMode::ON) do |model|
  model.context do |context|
    puts context.generate("Once upon a time")
  end
end

Use Llama::LazyMode::OFF to always read complete tensors up front.

The convenience API accepts the same setting:

response = Llama.generate(
  "/path/to/model.gguf",
  "Once upon a time",
  lazy_mode: Llama::LazyMode::ON
)

Resource Lifetime

Native-backed objects provide an idempotent free method. Prefer the block APIs shown above for deterministic cleanup. For longer-lived resources, call free in an ensure block and release dependencies before their owners:

model = Llama::Model.new("/path/to/model.gguf")
context = model.context

begin
  puts context.generate("Once upon a time")
ensure
  context.free
  model.free
end

Release samplers and adapters before contexts, and contexts before models. Samplers added to a SamplerChain are released with the chain.

Backend Lifetime

Llama.init is called automatically when a model or context is created, so most applications do not need to call it manually.

Llama.uninit is optional and usually not needed. It is intended only for controlled teardown after all Llama::Model and Llama::Context instances have been finalized. Calling it while models or contexts are still alive raises an error, because their finalizers may still need the llama.cpp backend.

Saved State Compatibility

llama.cpp b10809 updates the session and sequence-state file formats. Session or state files written by b10566 are not guaranteed to load with this version; recreate them after upgrading.

Advanced Sampling

require "llama"

Llama::Model.open("/path/to/model.gguf") do |model|
  model.context do |context|
    Llama::SamplerChain.open do |chain|
      chain.add(Llama::Sampler::TopK.new(40))
      chain.add(Llama::Sampler::MinP.new(0.05, 1))
      chain.add(Llama::Sampler::Temp.new(0.8))
      chain.add(Llama::Sampler::Dist.new(42))

      result = context.generate_with_sampler("Write a short poem about AI:", chain, 150)
      puts result
    end
  end
end

Chat Conversations

require "llama"

Llama::Model.open("/path/to/model.gguf") do |model|
  model.chat(system: "You are a helpful assistant.") do |chat|
    result = chat.ask("Hello, who are you?") do |chunk|
      print chunk.text
    end
    puts "\n#{result.finish_reason}"
  end
end

Chat history is committed transactionally. A cancelled turn is not committed unless commit_partial: true is requested. b10809 recognizes predefined chat template shapes; it is not a general Jinja evaluator. Pass a recognized template explicitly when the model does not provide one. Chat#clear removes every message, including the initial system message. A second turn, clear, or close during an active turn raises BusyError.

Embeddings

require "llama"

Llama::Model.open("/path/to/model.gguf") do |model|
  model.embedder(pooling: Llama::Pooling::Mean) do |embedder|
    vector = embedder.embed("Hello, world!", normalize: true)
    vectors = embedder.embed_all(["one", "two"], normalize: true)
    puts "Embedding dimension: #{vector.size}"
  end
end

Embedder owns a dedicated embedding context, copies native vectors before the next call, and preserves input order in native multi-sequence batches.

Utilities

System Info

info = Llama.runtime_info
puts "llama.cr #{info.wrapper_version} expects #{info.expected_build}"
puts "loaded llama.cpp #{info.reported_version} with #{info.backend_count} backends"
puts info.system_info

Tokenization Utility

Llama::Model.open("/path/to/model.gguf", vocab_only: true) do |model|
  puts Llama.tokenize_and_format(model.vocab, "Hello, world!", ids_only: true)
end

Examples

The examples directory contains sample code demonstrating various features:

API Documentation

See kojix2.github.io/llama.cr for full API docs.

API Layers

Llama.check_compatibility! can explicitly check the version reported by the library. It is not enforced during initialization, allowing other llama.cpp versions to be tested at the user's own risk. Stable packages of the supported release report 0.4.0, while the official b10809 release archives report 0.4.0-dev. The C API does not report the exact build number, so package pins and ABI checks are still needed when strict compatibility is required.

Custom Llama.log_set callbacks are experimental. On the pinned b10809 build, model loading and multithreaded decode callbacks were observed on the calling thread. Callback exceptions are contained at the C boundary and can be retrieved with Llama.take_log_callback_error.

Core Classes

Samplers

Development

See DEVELOPMENT.md for development guidelines.

This software is primarily created through AI-generated code.

Do you need commit rights?

Contributing

  1. Fork it (https://github.com/kojix2/llama.cr/fork)
  2. Create your feature branch (git checkout -b my-new-feature)
  3. Commit your changes (git commit -am 'Add some feature')
  4. Push to the branch (git push origin my-new-feature)
  5. Create a new Pull Request

License

This project is available under the MIT License. See the LICENSE file for more info.