Useful for comparing segmentation strategies before spending any API budget.
Arguments
- chunks
A gr_chunks object.
Value
A one-row data frame: method, n, total_tokens, min, median,
mean, max, over_cap. method reports any fallback that occurred.
total_tokens exceeds the document's own token count when overlap is on;
that difference is the duplication overlap buys you.
See also
Other segmentation functions:
gr_chunks,
gr_register_segmenter(),
gr_segment(),
gr_segment_spec(),
gr_segmenters(),
new_chunks()
Examples
doc <- gr_ingest(readgpt_example())
#> Using cached ingestion for this document + settings.
# What overlap actually costs, before any model call.
do.call(rbind, lapply(c(0, 30, 60), function(ov)
gr_chunk_stats(gr_segment(doc, list(method = "sentence", max_tokens = 120,
overlap_tokens = ov)))))
#> Segmenting with 'sentence' (cap 120 tokens, overlap 0).
#> Segmenting with 'sentence' (cap 120 tokens, overlap 30).
#> Segmenting with 'sentence' (cap 120 tokens, overlap 60).
#> method n total_tokens min median mean max over_cap
#> 1 sentence 6 532 47 92.5 88.7 106 0
#> 2 sentence 7 666 74 99.0 95.1 106 0
#> 3 sentence 9 860 79 94.0 95.6 109 0