module Llama
Defined in:
llama.crllama/adapter_lora.cr
llama/adapter_lora/error.cr
llama/batch.cr
llama/batch/error.cr
llama/cancellation.cr
llama/chat.cr
llama/context.cr
llama/context/error.cr
llama/embedder.cr
llama/error.cr
llama/generator.cr
llama/lib_llama.cr
llama/memory.cr
llama/memory/error.cr
llama/model.cr
llama/model/error.cr
llama/native_resource.cr
llama/options.cr
llama/result.cr
llama/runtime_info.cr
llama/sampler.cr
llama/sampler/adaptive_p.cr
llama/sampler/base.cr
llama/sampler/dist.cr
llama/sampler/error.cr
llama/sampler/grammar.cr
llama/sampler/grammar_lazy_patterns.cr
llama/sampler/greedy.cr
llama/sampler/infill.cr
llama/sampler/min_p.cr
llama/sampler/mirostat.cr
llama/sampler/mirostat_v2.cr
llama/sampler/penalties.cr
llama/sampler/temp.cr
llama/sampler/temp_ext.cr
llama/sampler/top_k.cr
llama/sampler/top_n_sigma.cr
llama/sampler/top_p.cr
llama/sampler/typical.cr
llama/sampler/xtc.cr
llama/sampler_chain.cr
llama/sampling_plan.cr
llama/session.cr
llama/session_snapshot.cr
llama/state.cr
llama/state/error.cr
llama/stop_detector.cr
llama/streaming_decoder.cr
llama/vocab.cr
Constant Summary
-
DEFAULT_SEED =
LibLlama::LLAMA_DEFAULT_SEED -
==== Native constants (wrapped for user convenience) ====
-
FILE_MAGIC_GGLA =
LibLlama::LLAMA_FILE_MAGIC_GGLA -
FILE_MAGIC_GGSN =
LibLlama::LLAMA_FILE_MAGIC_GGSN -
FILE_MAGIC_GGSQ =
LibLlama::LLAMA_FILE_MAGIC_GGSQ -
LLAMA_CPP_BUILD =
begin if match = VERSION.match(/^0\.(\d+)\.\d+$/) match[1] else VERSION end end -
LLAMA_CPP_COMPATIBLE_VERSION =
"b#{LLAMA_CPP_BUILD}" -
LOG_LEVEL_DEBUG =
0 -
Log level constants (from llama.cpp / ggml)
-
LOG_LEVEL_ERROR =
3 -
LOG_LEVEL_INFO =
1 -
LOG_LEVEL_NONE =
4 -
LOG_LEVEL_WARNING =
2 -
SESSION_MAGIC =
LibLlama::LLAMA_SESSION_MAGIC -
SESSION_VERSION =
LibLlama::LLAMA_SESSION_VERSION -
TOKEN_NULL =
LibLlama::LLAMA_TOKEN_NULL -
VERSION =
{{ (`shards version /srv/crystaldoc.info/github-kojix2-llama.cr-main/src`).chomp.stringify }}
Class Method Summary
-
.apply_chat_template(template : String | Nil, messages : Array(ChatMessage), add_assistant : Bool = true) : String
Applies a chat template to a list of messages
-
.builtin_chat_templates : Array(String)
Gets the list of built-in chat templates
-
.check_compatibility! : Nil
Explicitly verifies the loaded llama.cpp release against the version used to build these bindings.
-
.complete(model_path : String, prompt : String, options : GenerationOptions = GenerationOptions.new, &block : GenerationChunk -> ) : Generation
Performs a typed one-shot completion, optionally streaming safe text chunks.
- .complete(model_path : String, prompt : String, options : GenerationOptions = GenerationOptions.new) : Generation
- .error_message(code : Int32) : String
- .format_error(message : String, code : Int32 | Nil = nil, context : String | Nil = nil) : String
-
.generate(model_path : String, prompt : String, max_tokens : Int32 = 128, temperature : Float32 = 0.8, *, n_gpu_layers : Int32 = 0, offload_kqv : Bool = false, op_offload : Bool = false, lazy_mode : LazyMode = LazyMode::AUTO) : String
Generates text from a prompt using a model
-
.gpu_offload_supported? : Bool
Returns whether the current llama.cpp runtime supports GPU offloading.
-
.init
Thread-safe, idempotent initialization of the llama.cpp backend.
-
.llama_cpp_version : String
Returns the semantic version reported by the loaded llama.cpp library.
-
.log_level
Get the current log level
-
.log_level=(level : Int32)
Set the log level
-
.log_set(&block : Int32, String -> )
Set a custom log callback.
-
.max_parallel_sequences : Int64
Returns the maximum number of parallel sequences supported by backend This is a thin wrapper around LibLlama.llama_max_parallel_sequences.
-
.measure_ms(&)
Measures elapsed time in milliseconds for a block using llama.cpp's clock.
-
.mlock_supported? : Bool
Returns whether this llama.cpp build supports locking model data in memory.
-
.mmap_supported? : Bool
Returns whether this llama.cpp build supports memory-mapped model loading.
-
.process_escapes(text : String) : String
Process escape sequences in a string
-
.rpc_supported? : Bool
Returns whether the current llama.cpp runtime supports the RPC backend.
-
.runtime_info : RuntimeInfo
Collects version, backend, and feature diagnostics for the loaded runtime.
-
.system_info : String
Returns the llama.cpp system information
-
.take_log_callback_error : Exception | Nil
Returns and clears the first exception raised by the current log callback.
-
.time_ms : Int64
Returns the current time in milliseconds since the Unix epoch (llama.cpp compatible).
-
.time_us : Int64
Returns the current time in microseconds since the Unix epoch (llama.cpp compatible).
-
.tokenize_and_format(vocab : Vocab, text : String, add_bos : Bool = true, parse_special : Bool = true, ids_only : Bool = false) : String
Tokenize text and return formatted output
-
.uninit
Thread-safe, idempotent finalization of the llama.cpp backend.
Class Method Detail
Applies a chat template to a list of messages
Parameters:
- template: The template string (nil to use model's default)
- messages: Array of chat messages
- add_assistant: Whether to end with an assistant message prefix
Returns:
- The formatted prompt string
Raises:
- TemplateError if the template is unsupported or formatting fails
Gets the list of built-in chat templates
Returns:
- Array of template names
Explicitly verifies the loaded llama.cpp release against the version used to build these bindings. This check is opt-in so users can test other llama.cpp versions at their own risk.
Performs a typed one-shot completion, optionally streaming safe text chunks.
Generates text from a prompt using a model
This is a convenience method that loads a model, creates a context, and generates text in a single call.
response = Llama.generate(
"/path/to/model.gguf",
"Once upon a time",
max_tokens: 100,
temperature: 0.7
)
puts response
Parameters:
- model_path: Path to the model file (.gguf format)
- prompt: The input prompt
- max_tokens: Maximum number of tokens to generate (must be positive)
- temperature: Sampling temperature (0.0 = greedy, 1.0 = more random)
- n_gpu_layers: Number of model layers to offload to GPU (default: 0; negative = all layers)
- offload_kqv: Whether to offload KQV operations, including the KV cache (default: false)
- op_offload: Whether to offload host tensor operations to device (default: false)
- lazy_mode: Controls on-demand loading of eligible model tensors (default: LazyMode::AUTO)
Returns:
- The generated text
Raises:
- ArgumentError if parameters are invalid
- Llama::Model::Error if model loading fails
- Llama::Context::Error if text generation fails
Returns whether the current llama.cpp runtime supports GPU offloading.
Available dynamic backends are loaded by llama.cpp as needed.
Thread-safe, idempotent initialization of the llama.cpp backend. You do not need to call this manually in most cases.
Returns the semantic version reported by the loaded llama.cpp library. The exact build number is not exposed by llama.cpp's C API.
Set the log level
Parameters:
- level : Int32 - log level (0=DEBUG, 1=INFO, 2=WARNING, 3=ERROR, 4=NONE)
Example: Llama.log_level = Llama::LOG_LEVEL_ERROR # Only show errors Llama.log_level = Llama::LOG_LEVEL_NONE # Disable all logging
Set a custom log callback.
This bridge is experimental because callback thread behavior is controlled
by llama.cpp. The pinned b10809 build has been observed to invoke logging on
the calling thread, including during multithreaded decode. Exceptions are
caught at the native boundary and can be retrieved with
.take_log_callback_error.
The block receives:
- level : Int32 - log level (0=DEBUG, 1=INFO, 2=WARNING, 3=ERROR)
- message : String - log message
Example: Llama.log_set do |level, message| if level >= Llama::LOG_LEVEL_ERROR STDERR.print message end end
Returns the maximum number of parallel sequences supported by backend This is a thin wrapper around LibLlama.llama_max_parallel_sequences.
Measures elapsed time in milliseconds for a block using llama.cpp's clock.
elapsed = Llama.measure_ms do
# ... code to measure ...
end
puts "Elapsed: #{elapsed} ms"
Returns:
- Float64: elapsed milliseconds
Returns whether this llama.cpp build supports locking model data in memory.
Returns whether this llama.cpp build supports memory-mapped model loading.
Process escape sequences in a string
This method processes common escape sequences like \n, \t, etc. in a string, converting them to their actual character representations.
text = Llama.process_escapes("Hello\\nWorld")
puts text # Prints "Hello" and "World" on separate lines
Parameters:
- text: The input string containing escape sequences
Returns:
- A new string with escape sequences processed
Returns whether the current llama.cpp runtime supports the RPC backend.
Available dynamic backends are loaded by llama.cpp as needed.
Collects version, backend, and feature diagnostics for the loaded runtime.
Returns the llama.cpp system information
This method provides information about the llama.cpp build, including BLAS configuration, CPU features, and GPU support.
info = Llama.system_info
puts info
Returns:
- A string containing system information
Returns and clears the first exception raised by the current log callback.
Returns the current time in milliseconds since the Unix epoch (llama.cpp compatible).
t0 = Llama.time_ms
# ... some processing ...
t1 = Llama.time_ms
elapsed = t1 - t0
puts "Elapsed: #{elapsed} ms"
Returns:
- Int64: milliseconds since epoch
Returns the current time in microseconds since the Unix epoch (llama.cpp compatible).
This is a high-level wrapper for LibLlama.llama_time_us.
t0 = Llama.time_us
# ... some processing ...
t1 = Llama.time_us
elapsed_ms = (t1 - t0) / 1000.0
puts "Elapsed: #{elapsed_ms} ms"
Returns:
- Int64: microseconds since epoch
Tokenize text and return formatted output
This is a convenience method that tokenizes text and returns a formatted string representation of the tokens.
model = Llama::Model.new("/path/to/model.gguf")
result = Llama.tokenize_and_format(model.vocab, "Hello, world!", ids_only: true)
puts result # Prints "[1, 2, 3, ...]"
Parameters:
- vocab: The vocabulary to use for tokenization
- text: The text to tokenize
- add_bos: Whether to add BOS token (default: true)
- parse_special: Whether to parse special tokens (default: true)
- ids_only: Whether to return only token IDs (default: false)
Returns:
- A formatted string representation of the tokens
Thread-safe, idempotent finalization of the llama.cpp backend. Call this if you want to explicitly release all backend resources before program exit. All Model and Context instances must be released before calling this method.