---
title: "Guide: Adaptive Pairing"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Guide: Adaptive Pairing}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>"
)

cmdstan_available <- identical(Sys.getenv("PAIRWISELLM_RUN_CMDSTAN_VIGNETTES"), "true") &&
  requireNamespace("cmdstanr", quietly = TRUE) &&
  tryCatch(
    !is.null(cmdstanr::cmdstan_version()),
    error = function(e) FALSE
  )

knitr::opts_chunk$set(eval = cmdstan_available)
```

## Scope

This article shows how to rank **one set of items** with [adaptive_rank()]. The wrapper reads and
validates the items, starts or resumes a run, requests comparisons from a judge, performs Bayesian
refits, and returns the ranking and audit logs. Most users should start with this wrapper rather
than assembling the lower-level state and runner functions themselves. Executable CmdStan chunks
are disabled during ordinary package builds. Set `PAIRWISELLM_RUN_CMDSTAN_VIGNETTES=true` to opt in
when rendering this source locally.

## Choose a pairing strategy

Use exhaustive pairing when the set is small enough to evaluate every one of the
`choose(N, 2)` unordered pairs and complete coverage is important. Use a random sample when you
want a simple fixed-budget design or a baseline for comparison. Use adaptive selection when each
judgment is costly and you want the next comparison to respond to evidence already collected.

The runtime accepts `N >= 2` unique items. A range such as 30--2,000 items is operating guidance,
not an enforced limit. Very small sets may exhaust all eligible candidates before a Bayesian
refit, while large studies should be planned around provider cost, refit time, and storage.

## What happens during a run?

Within-set ranking has three phases.

1. **Phase 1: warm start.** A seeded shuffled chain schedules `N - 1` valid comparisons, producing
   a connected comparison graph.
2. **Phase 2: normal adaptation.** TrueSkill is updated after every committed comparison and guides
   the next pair. The controller balances coverage, long-, mid-, and local-range comparisons,
   repeated-pair limits, and fallbacks.
3. **Phase 3: near stop.** After a Bayesian refit passes diagnostics and its EAP reliability is
   within 0.05 of the stopping threshold, later refits use the stricter near-stop ESS requirement.
   Pair-selection behavior does not otherwise change merely because Phase 3 was entered.

TrueSkill is the fast, step-by-step selection model. Bayesian Bradley--Terry--Luce (BTL) refits are
slower and intermittent. They provide posterior item estimates and uncertainty, convergence
diagnostics, stability checks, and stopping decisions. Once an accepted posterior exists, its win
probabilities also gate eligibility for long-range comparisons; BTL is therefore not exclusively
a final reporting model.

```text
items -> warm start -> select one pair -> judge -> TrueSkill update
                              ^                    |
                              |                    v
                              +---- continue <- periodic BTL refit -> stop check
```

Adaptive selection is intended to direct effort toward useful comparisons, but it does not
guarantee a particular ranking, reliability, cost reduction, or improvement over random pairing.
Those outcomes depend on the items, judge, budget, and model assumptions.

## Install and check CmdStan

Bayesian refits require the suggested `cmdstanr` package, a working C++ toolchain, and CmdStan.
This is a one-time machine setup; it does not require an LLM credential.

```{r cmdstan-install, eval=FALSE}
install.packages(
  "cmdstanr",
  repos = c("https://stan-dev.r-universe.dev", getOption("repos"))
)
cmdstanr::check_cmdstan_toolchain(fix = TRUE)
cmdstanr::install_cmdstan()
```

Before starting a study, confirm that CmdStan can be found:

```{r cmdstan-check, eval=TRUE}
cmdstan_available

if (cmdstan_available) {
  cmdstanr::cmdstan_version()
}
```

The executable ranking chunks below run when CmdStan is available. The deterministic judge is
local, so the example never makes a network request or incurs provider charges.

## A complete offline run with `adaptive_rank()`

Start from a clean R session and load the package data. The `ID` column already contains unique,
non-missing identifiers. We retain `quality_score` only to simulate a judge.

```{r example-setup, eval=TRUE}
library(pairwiseLLM)

data("example_writing_samples", package = "pairwiseLLM")
samples <- example_writing_samples[, c("ID", "text", "quality_score")]

trait <- trait_description("overall_quality")
prompt_template <- set_prompt_template()

deterministic_judge <- function(A, B, state, ...) {
  a_wins <- A$quality_score[[1]] >= B$quality_score[[1]]
  list(
    is_valid = TRUE,
    Y = as.integer(a_wins),
    invalid_reason = NA_character_
  )
}
```

A custom judge receives one-row item tables `A` and `B` plus the current state. A valid response
must contain `is_valid = TRUE` and `Y` equal to `1` when A wins or `0` when B wins. The simulated
judge uses a fixture score so its choices are deterministic. Because it bypasses an LLM, the trait
and prompt are not consulted until the live example later in this article.

The minimum useful wrapper call supplies data, its ID and text columns, a judge, and a step budget.
We also choose a session directory so progress is persisted. With 20 items, the implemented
default `refit_pairs_target` is 20 committed comparisons: `ceiling(N / 2)` clamped to
`[20, 5000]`. A budget of 22 attempted steps leaves room for the warm start and produces one refit
under this deterministic fixture.

```{r adaptive-run}
session_dir <- tempfile("pairwisellm-adaptive-")

out <- adaptive_rank(
  data = samples,
  id_col = "ID",
  text_col = "text",
  judge = deterministic_judge,
  n_steps = 22L,
  session_dir = session_dir,
  persist_item_log = TRUE,
  resume = FALSE,
  seed = 42L,
  progress = "none"
)
```

`n_steps` is a maximum number of attempted steps for that call, not a promise that every attempt
will commit and not a global stopping criterion. A run can return earlier after Bayesian stopping
or candidate starvation.

### Inspect the ranking and stopping state

The wrapper returns the final state and common reporting views together. `out$items` is sorted by
the current within-set rank; lower rank values are better.

```{r adaptive-results}
out$summary

out$items[, c(
  "item_id", "theta_raw_eap", "theta_raw_sd", "rank_raw", "degree"
)] |>
  head()

out$refits[, c(
  "refit_id", "total_pairs_done", "diagnostics_pass",
  "reliability_EAP", "stop_decision", "stop_reason"
)]
```

The latent `theta_raw_eap` scale is relative within this set. Its sign and absolute magnitude are
not an external score scale. Posterior standard deviations describe model uncertainty under the
fitted BTL model; they do not include every possible source of judge or sampling error.

`out$summary$last_stop_reason` remains missing when the wrapper merely reaches its requested step
budget. For within-set runs, terminal reasons include `"btl_converged"` and
`"candidate_starvation"`. Inspect the last refit row as well as the run summary: no single
diagnostic should be interpreted as proof that the ranking is correct.

### Inspect comparisons and audit logs

```{r adaptive-logs}
names(out$logs)

step_log <- adaptive_step_log(out$state)
step_log[, c(
  "step_id", "pair_id", "A_id", "B_id", "Y", "status",
  "round_stage", "fallback_used", "starvation_reason"
)] |>
  head()

summarize_adaptive(out$state)
summarize_refits(out$state, last_n = 1L, include_optional = FALSE)
summarize_items(out$state, top_n = 5L)

result_history <- adaptive_results_history(out$state)
head(result_history)
```

The step log has one row per attempt. A non-missing `pair_id` identifies a committed result. The
round log has one row per completed Bayesian refit, and the item log stores posterior item
summaries for each refit. `adaptive_results_history()` converts committed outcomes to the
three-column format accepted by `build_bt_data()`.

`adaptive_get_logs()` returns the step, round, link-stage, and item-step views together. Use
`adaptive_step_log()` to audit calls and transactional validity, `adaptive_round_log()` for refit
diagnostics and stopping, and `adaptive_item_log(stack = TRUE)` for change across refits. Empty
typed tables are expected before the corresponding event occurs; they are not evidence that the
accessor failed. `summarize_adaptive()` is a compact view, not a diagnostic recomputation.

## Lower-level lifecycle and offline judges

Most users should use `adaptive_rank()`. The lower-level lifecycle is useful when an application
needs to hold the state itself: create it with `adaptive_rank_start()`, advance it with
`adaptive_rank_run_live()`, persist it, and reconstruct it with `adaptive_rank_resume()` or
`load_adaptive_session()`. Do not modify state lists or canonical log columns by hand.

```{r lower-level-lifecycle}
low_level_items <- samples[1:5, ]
low_level_items$item_id <- as.character(low_level_items$ID)

low_level_state <- adaptive_rank_start(low_level_items, seed = 17L)
low_level_state <- adaptive_rank_run_live(
  state = low_level_state,
  judge = deterministic_judge,
  n_steps = 3L,
  btl_config = list(refit_pairs_target = 5000L),
  progress = "none"
)

summarize_adaptive(low_level_state)
```

A custom judge has the contract `judge(A, B, state, ...)`. `A` and `B` are one-row item tables.
Return `list(is_valid = TRUE, Y = 1L)` when A wins or `Y = 0L` when B wins. To simulate refusal,
timeout, or parsing failure, return `is_valid = FALSE`, `Y = NA_integer_`, and a stable
`invalid_reason`. Use a local score column, a fixed lookup table, or a separately seeded random
generator to make simulations reproducible. The package's internal scenario harness is test
infrastructure, not a public workflow function.

`make_adaptive_judge_llm()` builds the same judge interface around `llm_compare_pair()`. It is a
lower-level alternative to supplying backend/model arguments to `adaptive_rank()`, and its calls
are live, billable, and subject to the selected provider's failure and privacy behavior.

## Persistence, interruption, and resume

Supplying `session_dir` makes the wrapper persist the initial state, every completed refit,
terminal stops, the end of an ordinary call, and intermediate checkpoints. The default checkpoint
cadence is every 100 attempted steps; change it with `checkpoint_every_steps` when losing that much
work would be costly. Frequent checkpoints increase disk I/O.

An abrupt process termination can lose attempts since the most recent completed save. After an
interrupt, validate what is on disk rather than assuming the in-memory state was written.

```{r adaptive-persistence}
session_metadata <- validate_session_dir(session_dir)
session_metadata[c("schema_version", "package_version", "n_items")]

loaded_state <- load_adaptive_session(session_dir)
summarize_adaptive(loaded_state)

resumed <- adaptive_rank(
  data = samples,
  id_col = "ID",
  text_col = "text",
  judge = deterministic_judge,
  n_steps = 2L,
  session_dir = session_dir,
  persist_item_log = TRUE,
  resume = TRUE,
  progress = "none"
)

c(
  before = nrow(out$logs$step_log),
  after = nrow(resumed$logs$step_log)
)
```

Resume is intentionally strict. If saved artifacts are present but invalid, `adaptive_rank()`
aborts instead of silently creating a new run. The input IDs and their order must exactly match
the saved session. Set `resume = FALSE` only when you deliberately want a new state, and choose a
new or empty session directory to avoid mixing studies.

The wrapper handles ordinary saving. For an explicit snapshot outside a wrapper call, use the
lower-level persistence helpers:

```{r adaptive-manual-save}
snapshot_dir <- tempfile("pairwisellm-snapshot-")
save_adaptive_session(resumed$state, snapshot_dir, overwrite = TRUE)
validate_session_dir(snapshot_dir)
snapshot <- load_adaptive_session(snapshot_dir)
```

Setting `persist_item_log = TRUE` writes a separate item-log file for each refit. The canonical
state, step log, round log, and metadata are persisted regardless; the option is useful when
per-refit item histories must be inspected independently.

## Invalid judgments and failed calls

Invalid judge output is transactional: it creates an attempted step with no committed `pair_id`,
does not update TrueSkill or comparison history, and does not count toward the BTL refit cadence.
It still consumes one unit of `n_steps`. The LLM judge created internally by `adaptive_rank()`
maps provider errors to invalid results and preserves available status, error, token, and raw
response fields in the step log.

This small example deliberately rejects its first attempt, then accepts the next one:

```{r invalid-judgment, eval=TRUE}
invalid_once_judge <- function(A, B, state, ...) {
  if (nrow(state$step_log) == 0L) {
    return(list(
      is_valid = FALSE,
      Y = NA_integer_,
      invalid_reason = "simulated_parse_failure"
    ))
  }
  deterministic_judge(A, B, state)
}

invalid_demo <- adaptive_rank(
  data = samples[1:4, ],
  id_col = "ID",
  text_col = "text",
  judge = invalid_once_judge,
  n_steps = 2L,
  resume = FALSE,
  seed = 7L,
  progress = "none"
)

adaptive_step_log(invalid_demo$state)[, c(
  "step_id", "pair_id", "status", "judge_valid", "judge_invalid_reason"
)]
```

Investigate repeated invalid rows before increasing the budget. Common causes include a thrown
judge error, a response that does not name either presented ID, or a custom judge that returns a
missing or non-binary `Y`.

## A detailed live-LLM workflow

The following is a **live-API example and is not evaluated**. It requires an OpenAI key, network
access, and incurs provider charges. The request shape shown here was recorded as tested with
pairwiseLLM 1.3.1 on 2026-09-05. That is dated evidence about this configuration, not a promise
that the provider still offers the model. Check
[Backends and Tested Model Configurations](https://shmercer.github.io/pairwiseLLM/articles/model-compatibility.html)
and the provider catalog
before a long run.

Set `OPENAI_API_KEY` outside the script. Replace `real_samples` with a data frame or file containing
one row per item, a unique ID, and the text to compare. Preserve that input in its exact row order
for resume.

```{r live-openai, eval=FALSE, purl=TRUE}
library(pairwiseLLM)

stopifnot(nzchar(Sys.getenv("OPENAI_API_KEY")))

real_samples <- utils::read.csv(
  "writing-samples.csv",
  stringsAsFactors = FALSE
)
stopifnot(nrow(real_samples) >= 2L)
stopifnot(all(c("ID", "text") %in% names(real_samples)))
stopifnot(!anyNA(real_samples$ID), !anyDuplicated(real_samples$ID))
stopifnot(!anyNA(real_samples$text), all(nzchar(real_samples$text)))

live_trait <- trait_description("overall_quality")
live_prompt <- set_prompt_template()
live_session <- "adaptive-live-session"

live <- adaptive_rank(
  data = real_samples,
  id_col = "ID",
  text_col = "text",
  backend = "openai",
  model = "gpt-5.6-luna",
  endpoint = "responses",
  trait_name = live_trait$name,
  trait_description = live_trait$description,
  prompt_template = live_prompt,
  judge_args = list(reasoning = "none"),
  n_steps = 200L,
  session_dir = live_session,
  checkpoint_every_steps = 10L,
  persist_item_log = TRUE,
  resume = FALSE,
  seed = 20260904L,
  progress = "refits",
  save_outputs = TRUE
)
```

This one call validates and normalizes the data, constructs the LLM judge, initializes the
within-set controller, attempts at most 200 judgments, performs Bayesian refits when due, persists
the session, and returns reporting views. `n_steps` is a per-call ceiling, not a target number of
valid comparisons or a guarantee that the stopping criteria will pass. Use a new or empty
`session_dir` with `resume = FALSE`.

The call intentionally retains the package's inference and stopping defaults. Do not copy reduced
diagnostic thresholds from a test fixture into a real study merely to make it finish sooner.

### Inspect the live run

Inspect the ranking, diagnostics, and stopping state before reporting results:

```{r live-openai-inspection, eval=FALSE, purl=TRUE}
live$summary

live$items[, c(
  "item_id", "theta_raw_eap", "theta_raw_sd", "rank_raw", "degree"
)] |>
  head(10L)

live$refits[, c(
  "refit_id", "total_pairs_done", "diagnostics_pass",
  "reliability_EAP", "stop_decision", "stop_reason"
)] |>
  tail()
```

An empty `live$items` means no successful BTL refit is available yet. Reaching the requested step
budget does not establish convergence. Read the latest refit diagnostics and stop reason together.

The step log records each attempted provider call, including provenance, validity, errors, and
available token counts:

```{r live-openai-audit, eval=FALSE, purl=TRUE}
live_steps <- live$logs$step_log[, c(
  "step_id", "pair_id", "A_id", "B_id", "Y", "status",
  "judge_backend", "judge_model", "judge_endpoint", "judge_invalid_reason",
  "llm_status_code", "llm_error_message",
  "prompt_tokens", "completion_tokens", "total_tokens"
)]

tail(live_steps)

colSums(live_steps[, c(
  "prompt_tokens", "completion_tokens", "total_tokens"
)], na.rm = TRUE)

live_steps[live_steps$status == "invalid", ]
```

Invalid API, refusal, or parse responses consume attempted steps but do not commit comparisons,
update TrueSkill, or advance the BTL refit cadence. Investigate repeated invalid rows before adding
budget. `include_raw = TRUE` on `adaptive_rank()` also stores serialized raw responses for deeper
auditing, but those payloads can contain submitted text and increase storage. Enable it only under
an appropriate retention policy.

### Continue the same study

Resume with the same input IDs, row order, trait, prompt, model settings, and session directory.
Only `n_steps`, progress display, and checkpoint cadence should normally change between calls.

```{r live-openai-resume, eval=FALSE, purl=TRUE}
live <- adaptive_rank(
  data = real_samples,
  id_col = "ID",
  text_col = "text",
  backend = "openai",
  model = "gpt-5.6-luna",
  endpoint = "responses",
  trait_name = live_trait$name,
  trait_description = live_trait$description,
  prompt_template = live_prompt,
  judge_args = list(reasoning = "none"),
  n_steps = 100L,
  session_dir = live_session,
  checkpoint_every_steps = 10L,
  persist_item_log = TRUE,
  resume = TRUE,
  progress = "refits",
  save_outputs = TRUE
)
```

Resume validates saved schemas and aborts on incompatible artifacts or mismatched input IDs instead
of silently starting a new run. A sudden process termination can lose attempts since the last
completed refit or checkpoint, so choose `checkpoint_every_steps` according to how many paid calls
you can afford to repeat.

Each live comparison is usually the dominant monetary and wall-clock cost. Bayesian refits are the
dominant local compute cost and become more expensive as the item set and posterior model grow.
Cost also rises with invalid calls and additional steps needed to clear diagnostics or coverage
gaps. Use bounded increments, inspect the audit and refit logs between calls, and do not add budget
when candidate starvation or a persistent structural problem requires a design change instead.

## Troubleshooting

**CmdStan is missing or compilation fails.** Run `cmdstanr::check_cmdstan_toolchain()` and
`cmdstanr::cmdstan_version()`. Install or repair the compiler toolchain before retrying. An LLM
credential cannot substitute for CmdStan.

**IDs are missing or duplicated.** `id_col` must identify a non-missing, unique column. For the
shipped data it is `id_col = "ID"`. Resume also requires the same IDs in the same order as the
saved state.

**The judge output is malformed.** A custom judge must return `is_valid`, `Y`, and preferably
`invalid_reason`. For a valid result, `Y` must be exactly `0` or `1`. Review
`judge_invalid_reason`, `llm_status_code`, and `llm_error_message` in `adaptive_step_log()`.

**The run reports candidate starvation.** The selector applies its implemented within-stage
fallbacks before declaring a stage starved. A terminal `"candidate_starvation"` means no eligible
pair remained after those fallbacks. This can occur with very small or heavily repeated designs;
inspect `fallback_path`, `starvation_reason`, and committed pair counts rather than treating it as
Bayesian convergence.

**There is no item summary yet.** `out$items` is empty until the first successful BTL refit. Under
the default cadence, at least 20 new committed comparisons are needed. Invalid and starved
attempts do not advance that count.

## Related documentation

See [adaptive_rank()], [adaptive_step_log()], [adaptive_round_log()], [adaptive_item_log()],
[save_adaptive_session()], and [validate_session_dir()] for complete argument and schema details.
For the within-set algorithm and statistical rationale, see
[Design: Adaptive Pairing](https://shmercer.github.io/pairwiseLLM/articles/within-set-adaptive-design.html).

## Citation

> Mercer, S. H. (2026). *Guide: Adaptive pairing* [R package vignette]. Comprehensive R Archive
> Network. https://doi.org/10.32614/CRAN.package.pairwiseLLM
