Tagged “generation”
-
Context Window
The maximum amount of text, measured in tokens, that a language model can process in a single request — input and output together.
-
Reranking
A second scoring pass that reorders an already-retrieved candidate set using a slower, more accurate model that reads query and document together.