Reliability is not validity
A measure can repeat consistently without representing the quality you care about. It can also be valid in principle but noisy in a field setting. Keep the two questions separate.
Log the setup
Warm-up, device, operator, surface, instructions, time of day and athlete state are part of the test. Recording them makes a future difference interpretable.
Choose a smallest worthwhile change carefully
Do not borrow a threshold from another population or device without checking the method. The practical threshold belongs to the test you actually run.
Reliability begins before the first test
Write the protocol as if another practitioner had to reproduce it. Include warm-up, instructions, device, operator, surface, timing, footwear, time of day, rest and the rule for accepting a trial. Decide whether the outcome is a best value, average, median or change score. If the athlete learns the task between visits, practice belongs in the protocol. These details are not administrative decoration. They are part of the measurement and explain why a number may move.
Interpret error at the level of the decision
A test can show a lower score without a real change in the quality you care about. Compare repeated trials, calculate or consult the typical error for the same method when available and ask whether the observed difference is larger than ordinary variation. Do not borrow a smallest worthwhile change from a different device or population. Reliability is not validity, and a reliable measure can still miss the performance construct. The goal is to know what size of change your own protocol can detect.
A simple example
If three sprint trials are close on Monday but widely spread on Thursday, the Thursday result may say more about start consistency than athlete adaptation. Repeating the setup before changing training is often the best next step. Keep context beside the number and mark device or operator changes clearly. A precise dashboard built on a variable protocol creates false confidence. The source trail points back to the CMJ monitoring literature because repeated measurement is a practical question, not just a statistical label.
Report uncertainty in plain language
Instead of saying that a score improved, say that the observed change was larger than, similar to or smaller than the usual variation of this protocol when that estimate exists. Show the number of trials and note any invalid attempt. If the method has not been studied in the exact setting, state that the uncertainty is unknown rather than importing confidence from a different device. A coach can still use the test, but the decision should be reversible and supported by other observations. Reliability is not a demand for perfect stability. It is a reason to match the precision of the decision to the precision of the measurement.
Precision should match consequence
Use high precision only when the measurement supports it and the decision is important enough to justify the effort. For a low-risk adjustment, a broad trend may be sufficient. For a consequential decision, uncertainty must be explicit.
Document what would change your mind
Before testing, state the size and direction of change that would justify action, as well as the result that would lead you to repeat the protocol. This protects the interpretation from hindsight. A measurement becomes more trustworthy when the practitioner can explain both the decision and the decision not to decide.
Editorial methods guide; see CMJ monitoring meta-analysis, PubMed 27663764.
Read the editorial method for how we handle evidence, limits and practical interpretation.

