The sample size is decided by the test, not by convenience
How many participants you need follows from what you intend to do with them, so it cannot be answered before the analysis is chosen. Students who collect first and ask afterwards find out that the design they wanted needed twice what they gathered, and by then the willing participants are gone.
It is also the calculation a committee is most likely to ask you to justify in a meeting. Doing it in advance produces a number with a reason attached; doing it afterwards produces a number with an apology attached, and the difference is very visible from the other side of the table.
Where the hours actually go in a quantitative study
Almost never in running the test itself, which takes minutes. They go into the preparation before it and the interpretation after it, and both are consistently underestimated by people planning a doctoral timeline.
Knowing the distribution is what makes the stage plannable. The six lines below are what our appraisals keep finding, and only one of them is what most students picture when they think about statistics.
- Cleaning: missing values, impossible entries, and deciding what to do about both
- Checking that what the test assumes is actually true of your data
- Running the analysis, which is genuinely the quickest part of the whole stage
- Rebuilding output into tables that follow your manual rather than the software
- Working out what the result means for the question you originally asked
- Writing it in sentences that survive somebody asking why
Collection takes longer than anyone budgets
The stage that wrecks doctoral timelines is not the analysis, it is getting hold of the data. Recruitment through a workplace waits on a gatekeeper. A survey distributed to a professional list returns at a rate nobody predicts. Retrospective records need somebody with access to pull them.
Not one of those depends on you, and every one of them can be estimated before it starts. Ask what the realistic response rate has been for similar work in your setting, double whatever period you first thought of, and start recruiting the day approval lands rather than the week you feel ready. Where the approval itself is still outstanding, that clock has not started yet.
Numbers that nobody in the room can explain
The failure mode specific to statistical help is a results chapter that is technically correct and completely unowned. It happens when somebody hands over data, receives output, and pastes the output in without ever being told what was decided along the way.
It fails in exactly one place, which is the meeting. A committee asks why this test rather than another, what happened to the participants who dropped out, and whether the finding would survive a different specification. Those are ordinary questions and they are unanswerable by somebody who was not part of the decisions, which is why the decisions get explained here as they are made.
You have to be able to say it out loud
Everything here is built with you rather than delivered to you, and the reason is a defense meeting. The questions coming at you are why this procedure, why this many participants, what the study cannot show and what a second attempt would change, from readers who work with numbers professionally.
So the plan gets explained while it is being written, and the findings are discussed with you instead of delivered to you, which takes an extra hour and is the entire point. Where the meeting itself is what worries you, that is rehearsed separately. What is never done is fabricating data, adjusting it toward a preferred result, or describing something as collected when it was not. That is refused at the enquiry, and it is refused for reasons that outlive any deadline.
Send the question, before anything else
The research question as currently worded, the population you intend to reach and what your committee has said so far. That is enough to say whether the design supports the question, which is the only thing worth knowing at this stage.
An analysis plan a committee will accept
Which test answers which question, what has to be true for that test to hold, how many participants it needs, and what you will do when an assumption fails. Written before collection, so nobody can say it was chosen to suit the result.
Results in sentences you can defend
Output rebuilt into properly formatted tables, the figures that belong in a paper separated from the diagnostics that do not, and each result written as a sentence you could say out loud in a room and be questioned on.
Questions the desk is asked.
When should the analysis be decided?
Before collection, always. The test you intend to run determines how many participants you need and what has to be true of the data, so choosing afterwards means discovering that the design you wanted needed twice what you gathered. It also protects you from any suggestion that the analysis was chosen to suit the result.
How many participants do I need?
It depends entirely on what you intend to do with them, which is why the question cannot be answered before the analysis is chosen. Calculated in advance it produces a number with a justification attached, which is exactly what a committee will ask for. Calculated afterwards it produces a number with an excuse attached.
What if an assumption fails?
Say so in the write-up, then choose. You can move to a procedure that does not depend on the assumption, you can adjust the data and declare exactly what you adjusted, or you can carry on and record the violation among your limitations. Each of those is defensible in a meeting. Saying nothing is the only route that is not.
How long does data collection actually take?
Longer than the plan says, and it is the stage that ruins doctoral timelines rather than the analysis. Recruitment waits on gatekeepers, survey response rates disappoint, and records need somebody with access to pull them. Ask what similar work in your setting achieved, then double your first estimate.
Can you just run it and send the output back?
The analysis yes, a set of numbers with no explanation no. A committee will ask why this test, what happened to the dropouts and whether the finding survives a different specification. Those are ordinary questions and they are unanswerable by somebody who was not part of the decisions.