How it was done

Methodology

How we selected, tagged, validated, and quality-checked the accounts that make up this analysis, and where the numbers come from.

The source: AHRERC digitised accounts

The accounts analyzed here come from the digitised portion of the Alister Hardy Research Centre (AHRERC) collection, the subset of the full 6,700+ account archive that has been made available in digital form. We accessed these through the RERC for research purposes.

Not every account in the digitised collection was included in this analysis. We applied selection criteria to focus on accounts that met our definition of first-hand individual experience accounts.

Selection criteria

We included accounts that:

  • Were written in the first person by the person who had the experience
  • Described a specific experience (not general spiritual views or commentary on the project)
  • Were of sufficient length to be meaningfully tagged (at least two paragraphs)
  • Described the experience of a single individual (not family reports or secondhand accounts)

We excluded accounts that were letters commenting on Hardy's project rather than describing personal experience, accounts that clearly reported someone else's experience, and accounts too brief to tag reliably. The 2,080 accounts in this analysis represent those that passed all selection criteria and quality review.

Tagging pipeline

Each account was processed through an automated tagging pipeline followed by a cross-validation stage. The pipeline ran in four stages:

  1. Initial tagging (Claude claude-opus-4-8)Each account was sent to Claude claude-opus-4-8 with a structured prompt specifying the six tagging dimensions and their permitted values. The model returned a JSON object with values for all dimensions. Accounts were processed in batches with rate limiting and error handling.
  2. GPT-4o cross-validationEach tagged account was independently re-evaluated by GPT-4o, which checked the initial tags and flagged any dimensions where it disagreed. Disputed dimensions were sent to an adjudication round.
  3. AdjudicationOn disputed dimensions, a second Claude pass received both the original and the GPT-4o assessment and was asked to choose the best answer with reasoning. The adjudicated answer became the final tag for that dimension.
  4. Human quality reviewA random sample of accounts at each quality tier was reviewed manually. Accounts assigned quality: low were inspected before deletion decisions were made. All deletion decisions were made by a human reviewer, not the pipeline.

Tagging schema

Each account was tagged across six dimensions:

DimensionValuesNotes
experience_categories 17 terms (multi-select) What type of experience occurred. Multiple categories could apply to one account.
phenomenological_qualities 9 terms (multi-select) Inner qualities of the experience, based on William James's framework and extensions. Multiple qualities could apply.
trigger 11 terms (multi-select) What was happening when the experience began. Multiple triggers could apply.
tradition_framing 7 terms (single value) The religious or cultural framework in which the person understood and described their experience.
lasting_effects 8 terms (multi-select) What changed permanently as a result of the experience, as described by the account author.
veridical_claim true / false Whether the account contained a claim of anomalous knowledge, information obtained during the experience that could not have been known by ordinary means.

Quality assessment

Each account was also assigned an interview quality rating (high / medium / low) reflecting the richness and specificity of the experience description. This rating was used during the selection and review process but is not displayed on the website, all 2,080 accounts in the analysis met a minimum threshold of medium or high quality.

Statistics and percentages

All percentages shown on this website are computed from the 2,080 accounts that passed the full pipeline. For multi-select dimensions (experience categories, qualities, triggers, lasting effects), percentages reflect the proportion of accounts in which each term appeared, they will therefore sum to more than 100%.

For tradition framing, two percentages are available: the percentage of all 2,080 accounts, and the percentage of the 2,063 accounts where a tradition was identifiable (17 accounts had insufficient framing language to assign a tradition).

All statistics on the website are generated programmatically from the underlying data and are updated automatically whenever the dataset changes. No statistics are hardcoded into the website's HTML.

Limitations and caveats

The archive has significant demographic skew: its respondents were predominantly British adults who saw Hardy's newspaper appeals in the 1960s–1980s, a period and context that heavily overrepresents Christian backgrounds and English-speaking adults. Non-Christian traditions are substantially underrepresented relative to their share of the world's religious population.

The accounts were self-selected by people who chose to write in, a significant sampling bias toward those who considered their experience worth reporting. People who had experiences they dismissed, forgot, or felt unable to describe are not represented.

Tagging is not infallible. The pipeline achieves high accuracy on most dimensions but makes errors, particularly in accounts that are ambiguous, metaphorical, or that span multiple categories with approximately equal weight. The two-model cross-validation significantly reduces systematic errors but cannot eliminate them. Where precision matters, go back to the primary accounts rather than relying solely on the aggregate data.