An envelope is opened. A word on the paper matches the word a spectator chose. The audience has seen a striking result. Whether it has seen evidence for extraordinary knowledge depends on events that may have received little attention: when the paper was written, who handled it, how the word was selected, and whether other predictions were available.
The envelope is a prop until those conditions become part of the record. Its presence does not supply the conditions by itself.
A useful test begins by converting a broad claim into something a person could get wrong. “I sense what people are thinking” leaves much undecided. “I can identify which of five specified symbols was independently selected, before anyone gives me feedback” names an answer, a time, and a set of alternatives. It is still only a claim, but now the result can be examined.
Agree on what the ability is supposed to do
Different claims require different conditions. Knowing what a nearby person sees is not the same as predicting a selection made later. Describing a concealed physical object is not the same as producing words that a listener finds meaningful. A test designed for one should not quietly be presented as a test of all three.
The person making the claim needs an opportunity to state its conditions beforehand. If distance, timing, an assistant, or a particular kind of target is said to matter, write that down. The evaluator can then decide whether the conditions allow an informative comparison. Cooperation does not require accepting a procedure in which ordinary information remains available.
This conversation should also settle the response format. A list of possibilities is different from one answer. A description interpreted after the target is revealed is different from an exact symbol named beforehand. Either may be studied, but the scoring rule must fit what was agreed. Changing from exact matches to generous resemblance after a disappointing result changes the question.
Make the chance baseline visible
Here is an illustrative test design, not a report of an experiment. Suppose the targets are circle, square, triangle, star, and cross. On each trial, an independent process selects one of them with equal probability, with the same five options available again on the next trial. The respondent records one answer before receiving any information about the target.
With independent uniform selection and no target information, each answer has a one-in-five chance of matching. In a fixed run of twenty trials, the expected number of matches is four. Expected does not mean guaranteed: chance runs can score above or below four.
For this example, suppose the evaluator chooses nine or more exact matches as a screening threshold before the run. Under the stated independent chance model, the probability of nine or more matches in twenty trials is about 1%. This is a binomial calculation, obtained by adding the probabilities for nine through twenty matches. No participant trials were run to produce it.
A score clearing that threshold would warrant examination under the chosen rules. It would not mean “there is a 99% probability that the ability is paranormal.” The calculation asks how often the score would arise under one specified chance model. It does not assign probabilities to every possible explanation, including information leakage or a faulty selection process.
Nor is twenty trials a general recommendation for a scientifically decisive test. A real investigation needs sample-size planning for the effect it seeks, scrutiny of its apparatus, and an agreed analysis. The small example makes the logic visible; it does not certify a research protocol.
Protect the target, the answer, and the scoring
In the proposed comparison, ordinary information must be unavailable before the answer is fixed. If someone who knows the target sits facing the respondent, expressions, speech, or timing become possible information paths. If the target sits behind a thin sheet of paper, the physical arrangement also deserves inspection. Confidence that nobody intends to help is different from evidence that they cannot inadvertently help.
The answer needs similar protection. It should be recorded in a form that cannot be quietly revised after a hint or after the target is revealed. The record should preserve the original answer, the target, the order, and any interruption. An evaluator should not have to reconstruct these from memory afterward.
For exact symbols, scoring is comparatively straightforward. For descriptions, independent scorers should not know which description is supposed to match which target while judging similarity. Otherwise knowledge of the desired pairing can influence the evaluation. If a scoring rule cannot be stated clearly enough for another person to apply, a striking agreement may conceal interpretive flexibility.
A rehearsal can check instructions and equipment, but its outcomes must stay distinct from the scored run. Discovering that the cards are marked or a screen reflects in a window would reveal a flaw in the apparatus. It would not be evidence for or against an extraordinary faculty. Repair the flaw, preserve the record of the change, and agree on the test that follows.
Count the unsuccessful trials, and stop where you said you would
Suppose the proposed twenty-trial run reaches eight matches. Continuing until a ninth match appears would make the observed success depend on a stopping rule different from the one used in the calculation. Reporting only the best run from many undisclosed runs creates another problem. The chance comparison must account for the opportunities that actually occurred.
Fix the run length, exclusions, and treatment of incomplete sessions beforehand. If a session stops because a participant wishes to leave, record it as incomplete. Consent should remain real even when an evaluator wants a finished score. Do not convert a withdrawal into a selectively discarded bad run and begin again until the results look favorable.
The same principle applies to changing the claim. If circle and square are missed but the respondent says their “energy” was similar, that is a new scoring proposal. It can be defined for a future comparison. It cannot retroactively turn the old misses into successes without changing the evidence.
A published result can also fail to travel
A historical research case makes replication concrete. In 2012, Stuart Ritchie, Richard Wiseman, and Christopher French published three preregistered attempts to reproduce a reported retroactive-recall effect. The earlier claim concerned whether words practiced after a memory test were already better recalled before that practice. The replications involved fifty participants each, used closely matched procedures with some localized vocabulary, and did not find support for the effect. Their word-coding procedure also kept scorers blind to the words’ later practice status. The paper’s methods, results, and discussion.
This case concerns that particular recall claim. It is not a complete survey of paranormal research, and it is not a verdict on every experience people describe. Its value here is methodological: a procedure can be specified, attempted independently, and produce results that fail to support the original interpretation.
Replication does more than add another impressive number. It asks whether a result survives a new team, new observations, and safeguards fixed before anyone knows the outcome. An explanation that depends on a single favorable arrangement has less support than one that continues to work under independent examination.
Keep a confirmation block out of the rehearsal
For the illustrative five-symbol design, an evaluator could reserve a second, independently selected twenty-trial block for confirmation. Its targets would remain unavailable while the procedure was adjusted. The same screening criterion would be fixed in advance, and both blocks would be reported separately, including failures.
If the first block prompted new instructions, new exclusions, or a new scoring rule, it would become development evidence. The untouched block would then test the revised proposal. Once that block too influenced changes, another fresh comparison would be needed. Calling repeatedly reused outcomes “independent confirmation” would conceal the tuning.
Even two unusual scores would leave the physical and procedural questions important. An information leak can repeat. A selection flaw can repeat. Independent review should try to find those counterexamples rather than treating the low chance probability as permission to stop asking how the test worked.
Let an unresolved result remain unresolved
Three statements can coexist: the performance was convincing, its method remains unknown to you, and the available record does not establish the proposed extraordinary explanation. None requires mocking the performer or the spectator.
A fair inquiry gives a claim somewhere definite to succeed or fail. If the conditions cannot be agreed, the claim remains untested under those conditions. If the record is incomplete, the conclusion stays limited. If an unusual result survives the planned comparison, it earns further investigation before a larger explanation is adopted.
The difference between a show and evidence is the work of making the result answerable to conditions the show did not have to satisfy. A sealed envelope can preserve a prediction. Only a preserved procedure can tell us what the prediction establishes.
Loading comments…