class CodeLocator

Defined in:

models/code_locator.cr

Constant Summary

DEFAULT_CONTENT_CACHE_BUDGET = (512_i64 * 1024) * 1024

Default content cache budget (bytes). Override via NOIR_CONTENT_CACHE_MAX_MB (value in megabytes). Set to 0 or the env NOIR_CONTENT_CACHE_DISABLE=true to disable caching entirely, in which case #content_for always returns nil and analyzers fall through to File.read.

MAX_CONTENT_CACHE_MB = Int64::MAX // (1024_i64 * 1024)

Megabyte figures at or above this overflow Int64 once scaled to bytes. Crystal checks integer arithmetic, so the multiply below raised rather than wrapping: NOIR_CONTENT_CACHE_MAX_MB=99999999999999 aborted the whole scan with an unhandled OverflowError stack trace before the first file was read. Saturate instead — a budget that large already means "cache everything", which is exactly what Int64::MAX gives.

Constructors

Instance Method Summary

Constructor Detail

def self.instance : CodeLocator #

[View source]
def self.new #

[View source]

Instance Method Detail

def all(key : String) : Array(String) #

[View source]
def base_relative(path : String) : String #

path relative to the scan base that owns it, /-separated and rooted with a leading /. With no registered bases (library callers, unit specs that drive a parser directly) the path is returned unchanged, which is exactly the pre-registry behaviour.


[View source]
def build_basename_index #

Build a basename => paths index from file_map.

The companion to #build_extension_index, for the "find the file(s) whose path ends with a/b/c.py" lookups that analyzers otherwise answer with Dir.glob("root/**/a/b/c.py"). That glob walks the whole tree from the scan root — descending into node_modules, .git and every other subtree the detector deliberately pruned — and costs one opendir/getdirentries pass per call. Django's ROOT_URLCONF resolution ran it once per settings.py, which on a 44k-file monorepo was ~73% of the entire analysis phase.

Keyed on basename because that is the selective part of those patterns: urls.py narrows 44k files to a handful, and the caller confirms the rest of the path with a cheap ends_with?.


[View source]
def build_extension_index #

Build extension index from file_map for fast lookups


[View source]
def clear(key : String) #

[View source]
def clear_all #

[View source]
def content_cache_stats : NamedTuple(bytes: Int64, files: Int32, skipped: Int32, budget: Int64) #

[View source]
def content_for(path : String) : String | Nil #

Returns cached file content or nil if the file was not cached (budget exhausted, caching disabled, or read after cache was cleared). Callers should fall back to File.read on nil.


[View source]
def expanded_file_map : Array(Tuple(String, String)) #

{original, File.expand_path(original)} for every file in file_map, built once and cached. File.expand_path is pure string normalization but non-trivial, and the monorepo helpers in FileHelper re-scan all_files once per base path per analyzer — without this the same path is expanded thousands of times (O(analyzers × bases × files)). Invalidated whenever file_map changes (push / clear).


[View source]
def expanded_path_for(path : String) : String #

O(1) path => File.expand_path(path) lookup for files registered in file_map. Analyzers call path_under_root?(file, base) inside base_paths.each { files.each { ... } } loops, so the same file would otherwise be re-expanded once per base (and File.expand_path of a relative path issues a getcwd). Unregistered paths fall back to a live expansion. Shares the lazy lifecycle / invalidation of #expanded_file_map.


[View source]
def files_by_basename(basename : String) : Array(String) #

Files whose basename is exactly basename (O(1) lookup).


[View source]
def files_by_extension(extension : String) : Array(String) #

Get files by extension using the index (O(1) lookup)


[View source]
def get(key : String) : String | Array(String) #

[View source]
def push(key : String, value : String) #

[View source]
def register_file(path : String, content : String) #

One-shot used by the detector's file reader: push the path into file_map and (budget permitting) cache the content so analyzers can skip the second File.read. Files whose content exceeds the remaining budget are still registered in file_map but not cached, and #content_for(path) returns nil for them — callers must keep a File.read fallback.


[View source]
def scan_base_paths : Array(String) #

[View source]
def scan_base_paths=(paths : Array(String)) #

The -b roots of the current scan, published once by NoirRunner.

Convention filters ("is this file under tests/?", "is this bundled output?") must run on the path relative to the scan base, never on the absolute path — otherwise a directory above the base decides the result and the same source tree reports different endpoints depending on where it is checked out. Analyzers get that through Analyzer#base_relative_path; the shared parser layer (Noir::JSRouteExtractor and friends) has no analyzer instance, so it reads the roots from here.


[View source]
def set(key : String, value : String) #

[View source]
def show_table #

[View source]