# Preference Ranking Instructions

**Version:** 1.0 · **domainXpert Annotation Standards**

---

## Overview

Preference ranking is the process of evaluating two or more model-generated responses and indicating which best satisfies the task requirements. Your expert judgment is the signal — this guide ensures that signal is consistent, calibrated, and useful.

---

## Core Principles

1. **Evaluate the response, not the style.** A well-formatted but factually weak response should rank below a plainly written but accurate one.
2. **Use your domain knowledge.** You were selected because you have expertise the model lacks. Apply it.
3. **Be consistent.** Apply the same standards across all pairs in a batch.
4. **Avoid length bias.** Longer responses are not inherently better.

---

## Scoring Dimensions

| Dimension | Weight | Description |
|---|---|---|
| Factual Accuracy | High | Are all claims correct and verifiable? |
| Completeness | Medium | Does the response address all parts of the prompt? |
| Reasoning Quality | High | Is the logic sound and well-structured? |
| Appropriate Uncertainty | Medium | Does the response hedge where appropriate? |
| Clarity | Low | Is the response easy to follow? |

---

## Step-by-Step Process

1. Read the prompt carefully. Identify what a correct, complete answer requires.
2. Read Response A in full before reading Response B.
3. Score each response independently on the dimensions above (1–5 scale).
4. Select your preference: A, B, or Tie (use Tie sparingly).
5. Write a brief justification (2–4 sentences) citing specific strengths or weaknesses.

---

## Common Errors to Avoid

- **Anchoring:** Don't let the first response set your expectations for the second.
- **Recency bias:** Don't favor the response you read most recently.
- **Sycophancy detection:** Flag responses that agree with the prompt's framing without independent verification.

---

*For questions, contact your project coordinator.*
