Methodology
How your diagnosis is built
This assessment is a diagnosis, not a scorecard. Its whole job is to find the earliest place your ministry loses momentum (the one thing worth fixing first) and to tell you plainly where not to spend your energy yet. This page walks through exactly how it works, step by step, in the order the diagnosis is assembled.
What this assessment is
Most church surveys hand you a wall of numbers and leave you to guess what matters. This one does the opposite. It treats your ministry as a journey people move through (from first-time guest to connected, mature, serving, generous, and multiplying disciple) and looks for the first point where that journey stalls.
That first stall is your primary constraint: the single place where effort will do the most good right now. Everything downstream of it is usually a symptom, not a cause. So the report is deliberately opinionated: it names one thing to work on, and just as importantly names the things to leave alone for now.
What we measure
We look at eight areas. Five are sequential stages, the journey itself, read in order:
- Guest Experiencethe front door
- Connectionbelonging
- Discipleshipgrowth
- Volunteeringserving
- Generositygiving
The other three are enablers, the foundations that sit under every stage:
- Governance & Accountabilitygates every stage
- Communicationgates the front half
- Structure & Systemsgates serving & growth
Each area is measured by a small set of questions answered on a 1 to 10 scale. Every question is anchored, with the 1, the middle, and the 10 each described in words, so a “7” means the same thing to everyone answering. Questions are also tagged as belief (“what you think is true”) or evidence (“what’s actually happening”). That distinction powers the blind-spot check later on.
How a score is formed
An area’s raw answers are averaged and rescaled to a familiar 0 to 100 range. But a simple average has a well-known flaw: some people rate everything harshly and some rate everything generously, so a score can say more about who answered than about the ministry itself.
To correct for that, we separate the area’s true level from each rater’s personal tendency, and report the level. Two churches with identical ministries but different personalities on the team end up with comparable scores. We also only count a person toward an area once they’ve answered all of that area’s questions; partial responses are set aside rather than allowed to skew the result.
The two headline numbers
The report leads with two numbers. Your health score is how you’re doing on average across all eight areas, your raw strength. Your real-world result is how well the whole chain actually moves people all the way through.
The real-world result is weighted heavily toward your weakest stage, on purpose: a chain is only as strong as its weakest link, and one blocked stage caps how many people reach the end no matter how strong the others are.
The distance between the two — the points lost to your weakest area — is the revealing part. A wide gap means you have real strength that isn’t translating into end-to-end flow: hidden drag. A narrow gap means what you’ve built is actually carrying people through.
Each of your eight areas is then read against a single bar we call the standard: 80 out of 100. Anything below it is real room to improve, and the dashboard counts against that bar — how many areas sit below the standard, and which three to work on first.
The colour on each area is a separate reading, and a deliberately gentler one: it describes how an area is holding up in its own right, not how far it is from the standard. So an area can be coloured as a strength and still be below the standard — that is not a contradiction. It is doing well compared with where churches usually are, and it still has room to grow.
The chain and its dependencies
Because the stages are a sequence, we read them in order. The first stage that falls below the line is your primary constraint: the one thing to fix first. Stages that break further down the chain are usually downstream symptoms, so the report tells you to leave them for now.
Enablers don’t get their own headline; instead they gate the fix. Weak governance can block progress everywhere; weak communication tends to hurt the front of the journey; weak structure tends to hurt serving and growth. If an enabler is holding a stage back, fixing the stage alone won’t stick.
We describe each dependency between two areas in plain words:
- Load-bearing: both ends are weak, and the relationship is actively costing you.
- Clear: the upstream area is strong, so it isn’t the explanation for what’s downstream.
- At risk: downstream looks fine but rests on a weak upstream, so you’re running on borrowed time.
- Solid: both ends are strong, so there’s nothing to flag.
Blind spots: belief vs evidence
Remember that each question is tagged as belief or evidence. When your belief runs well ahead of the evidence, that’s a blind spot: you feel better about an area than what’s actually happening supports. When the evidence runs ahead of belief, you may be underrating yourselves and missing a real strength.
A few areas are measured by belief only, so they carry no evidence cross-check: the report is explicit about that rather than inventing a comparison.
Agreement and confidence
Before we ask whether your team genuinely disagrees about an area, we first remove each person’s rating style, so one habitually harsh rater doesn’t look like conflict when there is none. What’s left is real disagreement, and we flag it when it’s meaningful, because a split team often signals an area worth a closer look.
We also lower our stated confidence when an area rests on very few responses. A verdict built on two answers is treated more cautiously than one built on twenty.
The role of AI
This matters, so we’ll be direct. Every score, constraint, and verdict is decided by a deterministic engine: fixed rules and math that produce the same answer every time from the same answers.
AI is used only to rephrase findings that are already decided into readable prose, and its wording is fact-checked back against the numbers. No model ever changes a score or a verdict. The diagnosis is the math; the AI just helps it read like a person wrote it.
Versioning and honesty
This methodology is versioned, and we treat its thresholds and benchmarks as tunable: they’re being refined as more churches take the assessment and more real data comes in. We’d rather tell you that openly than pretend the model is finished.
If you’d like to talk through your own results and turn them into a plan, the best next step is a conversation.
Talk through your results
Book a free call with the XP Gathering team, and we’ll walk through your diagnosis together and map the next few moves for your church. No cost, no pressure.
Book your free call →