Describe a reading configuration
Usage
gr_read_spec(
reader = "map_reduce",
model = NULL,
temperature = NULL,
max_answer_tokens = 1500L,
max_chunk_tokens = 700L,
max_summary_tokens = 500L,
top_k = 6L,
min_score = -Inf,
mmr = 1,
context_order = c("relevance", "document", "edges"),
rerank_candidates = 20L,
rerank_min_score = 4,
fan_in = 5L,
max_levels = 5L,
max_rounds = 4L,
preview_tokens = 1200L,
restate = c("auto", "always", "never"),
members = NULL,
cite = FALSE,
skim_model = NULL,
summary_model = NULL,
parallel = NULL,
delay_between_calls = 0,
on_overflow = c("warn", "error"),
...
)Arguments
- reader
Reader name; see
gr_readers().- model
Chat model id.
- temperature
Sampling temperature, or
NULLto omit the field. You do not need to null it yourself for reasoning models: it is dropped automatically for any model whose registry entry hassupports_temperature = FALSE, which includes the default model. Check withgr_model_info(model)$supports_temperature.- max_answer_tokens
Output cap for final answers and merges.
- max_chunk_tokens
Output cap for per-chunk calls.
- max_summary_tokens
Output cap for summarisation calls.
- top_k
For
retrieve,rerankanditerative: chunks to use.- min_score
For
retrieve: chunks scoring below this are dropped, but if that would leave nothing the single best chunk is used anyway.-Infdisables the filter. Not a cosine similarity when embeddings fall back to lexical vectors: the score is then a blend of cosine and BM25.- mmr
Diversity of selection, for
retrieveanditerative.1(the default) is plain top-k. Below 1, chunks are picked greedily bymmr * relevance - (1 - mmr) * similarity to what is already picked, so three chunks saying the same thing do not all get in and pay for each other. Costs nothing: the vectors are already computed.0.7is a reasonable place to start;0selects for novelty alone and will happily pick irrelevant chunks because they are different.- context_order
Where the selected chunks sit in the prompt.
"relevance"(default) is most relevant first;"document"restores the order they appear in the document, which reads better when chunks are consecutive;"edges"puts the strongest first and second-strongest last, burying the weakest in the middle, because transformers attend measurably better to the beginning and end of a long context than to its middle. Selection is unaffected. This decides only placement, and it applies toretrieveandrerank, the two readers that put several ranked chunks in one prompt.- rerank_candidates, rerank_min_score
For
rerank: how many chunks to score, and the score below which a chunk is discarded.- fan_in, max_levels
For
hierarchical: summaries combined per call, and the recursion depth cap.- max_rounds
For
iterative: retrieve-assess cycles.- preview_tokens
For
preview: the cap on the outline the planner sees. The outline is built from section labels, sizes and short excerpts (never the full text), and per-section excerpts shrink until the whole thing fits, so every section stays visible to the planner rather than the outline being truncated and some sections never being offered to it at all. The planner is an LLM call about a long document and so prone to exactly the degradation this package manages; keeping its input small is the mitigation.- restate
Whether to repeat the question before the excerpts as well as after them:
"auto"when the body is long enough to bury the first ask,"always", or"never". A setting rather than a rule, because whether it helps is a question about your corpus and your model. Pointgr_compare()at two recipes differing only in this and find out.- members
For
ensemble: the reader names to combine.- cite
Ask for chunk-level citations (
[chunk 3]) in the answer. Map those ids back to pages viaans$evidence. Forced off forhierarchical, which answers from summaries: summaries carry no[chunk N]ids, so asking for citations there asks the model to invent them.- skim_model, summary_model
Optional cheaper models for the per-chunk stages.
skim_modelis used byskim's extraction andrerank's relevance scoring;summary_modelbyhierarchical's summarisation.- parallel
Run per-chunk calls in parallel. Needs the
futureandfuture.applypackages; without them the run is sequential and says so.The trace is complete either way: workers keep their own and the parent absorbs them, in input order, so a parallel run reports the same calls, tokens and cost as the same run made sequentially. Two things do not cross the process boundary. The limits in
gr_options()are checked before a batch is sent and not inside it, since a worker cannot see what the others spend, so the pre-flight check, which runs in the parent, is what bounds a parallel run: its estimated calls againstmax_calls, and its worst case, every reply at its cap, againstmax_cost_usd. And a client that keeps its own log in a closure, such asgr_mock_client(), only sees the calls made in this process; ask the trace instead.- delay_between_calls
Seconds to sleep between sequential calls, for rate-limit shaping. Honoured by
map_reduce,refineandskim; the other readers do not sleep.- on_overflow
For
stuff:"warn"(truncate and say so) or"error".- ...
Extra fields for custom readers.
Out-of-range values
Numeric arguments are clamped into a usable range and the change is warned
about, never applied silently: top_k [1, 1e4], fan_in [2, 32],
max_levels [1, 12], max_rounds [1, 20], rerank_candidates [1, 1e4],
rerank_min_score [0, 10], preview_tokens [100, 1e5],
delay_between_calls [0, 600],
token caps [16, 1e6].
See also
gr_readers() for the available readers and their call costs,
gr_read(), gr_recipe()
Other reading functions:
gr_answer,
gr_read(),
gr_reader_signature(),
gr_readers(),
gr_register_reader(),
is_not_found(),
new_answer()
Examples
gr_read_spec("retrieve", top_k = 8)$top_k
#> [1] 8
# Out-of-range settings are corrected loudly, not quietly.
suppressWarnings(gr_read_spec("hierarchical", fan_in = 999)$fan_in)
#> [1] 32