24,428 courses · 2,504 curriculum guides Sponsored by eAgentic Software Sponsored by eAgentic Software

STA4222: Sample Survey Design

STA4222 — Sample Survey Design
← Course Modules
3 credit hours 45 contact hours Prerequisites: Varies more than in any other course here, and it changes the course's depth. UWF requires MAC2311 (Calculus I) OR STA2023 (Elements of Statistics) -- a very wide door. Other Florida institutions require a calculus-based probability sequence such as STA4321/STA4322. After STA2023 the course is formula-based; after calculus-based probability it can derive the estimators and their variances. Intending graduate study in statistics? Take the probability sequence first. v1.0

Course Description

STA4222 Sample Survey Design is the course that takes seriously a question the rest of the statistics curriculum assumes away: where did the data come from? Most statistical training begins with a sample already in hand and asks what can be inferred from it. This course begins earlier, with a population that needs to be measured, a budget that will not permit measuring all of it, and a decision about how to select the part that will be measured — because that decision, more than any subsequent analysis, determines whether the answer will be right.

The course is offered at approximately eight Florida institutions, including the University of West Florida, the University of Florida, Florida State University, the University of Central Florida, the University of North Florida, Florida Atlantic University and the University of South Florida.

The title varies across the state in a way that signals a real difference in emphasis. The University of Florida titles it Sample Survey Design, which is the statewide title. The University of West Florida titles it Sampling Theory and describes it as "a first course in sampling methods with application to survey sampling and field sampling", covering simple random, stratified, cluster, systematic and adaptive sampling and the estimators corresponding to each design. The subject is the same; the centre of gravity differs. A course titled design leans toward the practical construction of a survey — frames, questionnaires, nonresponse, weighting. A course titled theory leans toward the derivation of estimators and their variances. Most versions do both, and the distinction is one of proportion rather than content.

UWF's inclusion of field sampling and adaptive sampling is worth noticing, and is not universal. Adaptive sampling — where the selection of later units depends on what earlier units revealed — is the design used when the thing being measured is rare and clustered: a fish population, a contaminated area, an invasive species. It is unusual in an undergraduate survey course and it is a strong fit for Florida, where the state's environmental and fisheries agencies do exactly this kind of work.

The intellectual content of the course is a single sustained argument: a probability sample of a thousand people supports inferences that a self-selected sample of a million does not. Students arrive in an era of abundant data and generally believe the opposite. Working through the mathematics of design-based inference — why the randomisation itself is what licenses the confidence interval, and why a large sample from a defective frame is not merely less accurate but wrongly centred — is the durable thing the course teaches. It is also the thing that makes its graduates useful, because the failure it guards against is now the most common failure in applied data work.

Learning Outcomes

Required Outcomes

Optional Outcomes

Major Topics

Required Topics

Optional Topics

Resources & Tools

Career Pathways

This course sits at an unusual intersection: it is mathematically substantial enough to prepare a student for graduate study in statistics, and practical enough to be immediately employable. The specific skill — designing a data collection that supports a defensible inference — is scarce and is not supplied by a general data science education.

Florida employers with genuine demand for this skill include the Bureau of Economic and Business Research at the University of Florida, which produces the population estimates that drive state revenue sharing and is one of the few places in the state doing official-statistics work; the Florida Department of Health, which administers BRFSS and other surveillance systems; the Florida Fish and Wildlife Conservation Commission and the NOAA Southeast Fisheries Science Center in Miami, where fisheries sampling design is core business; water management districts across the state, which run environmental monitoring programmes; the Florida Department of Transportation; academic health centres and cancer institutes including Moffitt in Tampa and the University of Florida and University of Miami health systems, which employ biostatisticians; and the state's market research, healthcare analytics and insurance sectors. The federal statistical agencies maintain regional operations in the state, and Florida's size makes it a substantial site for national survey data collection.

Special Information

⚠ Prerequisites — the widest variation in this batch, and it changes the course

The University of West Florida requires MAC 2311 (Calculus I) or STA 2023 (Elements of Statistics). That is a genuinely wide door: it admits both a student who has completed a calculus sequence and a student whose entire statistical background is one general-education course. Other Florida institutions gate the course more narrowly, commonly requiring a calculus-based probability and statistics sequence such as STA 4321 and STA 4322, or at minimum a second statistics course.

The consequence is that the same SCNS number denotes courses of substantially different mathematical depth. Taken after STA 2023, the course is necessarily formula-based: estimators and their variances are presented, applied and interpreted, with derivation kept light. Taken after a calculus-based probability sequence, the same course can derive the estimators, prove unbiasedness, work through the variance algebra, and treat the Horvitz-Thompson framework properly. Both are legitimate and both are useful; they are not interchangeable preparation for graduate study.

Practical guidance. If you intend graduate work in statistics or biostatistics, take the calculus-based probability sequence before this course regardless of what your institution enforces, and expect to encounter the derivations again at greater depth. If you are taking it as an applied methods course for research work in another discipline, the lighter prerequisite version will serve you well — the design reasoning, which is the part that prevents real mistakes, is fully available without the measure-theoretic machinery. If you are transferring this course, note that a receiving graduate programme may look at which prerequisite chain you completed rather than at the course number alone.

Position in the curriculum

STA4222 is an upper-division elective in statistics and data science programmes, normally taken in the junior or senior year. It is one of the few statistics electives that is equally valuable to non-majors, and it is commonly taken by students in public health, environmental science, political science, marketing, agriculture and wildlife ecology. It complements regression (STA 4210 or equivalent), experimental design, and categorical data analysis: experimental design and survey sampling are the two halves of the question "how do I get data that will answer this?", and a student who has taken both is unusually well equipped.

Articulation and transfer

STA4222 carries the same SCNS number across Florida public institutions and SCNS equivalency governs transfer of the credit. As an upper-division course it does not appear in A.A. programmes and is taken after transfer to a four-year institution. The standard caution applies with more force than usual here because of the prerequisite variation described above: SCNS equivalency determines that the credit transfers, and the receiving department determines whether it satisfies a major requirement at the depth that programme expects. Keep the syllabus.

Course format and workload

Three credit hours, approximately 45 contact hours, taught as lecture with problem sets; some institutions add a computing component and some offer the course online. Assessment normally combines problem sets, computing assignments, examinations, and — in the versions that emphasise design — a project in which students design a survey, and sometimes execute a small one. Expect six to nine hours a week outside class. The problem sets are the course: the algebra of variance under different designs does not become intuitive by reading about it.

What makes this course different from the rest of the statistics curriculum

Worth stating plainly, because it surprises students: in most statistics courses the randomness comes from the population — individuals vary, and the sample inherits that variation. In design-based survey sampling, the randomness comes from the sampling procedure itself. The population values are treated as fixed, unknown constants; what varies from one hypothetical repetition to another is which units were selected. This is a genuine conceptual shift, and it is the reason the confidence interval means what it means. Students who try to map every result back onto the model-based framework they learned earlier find the course harder than it needs to be. It is worth pausing on the distinction early.

Why this course has become more valuable, not less

A reasonable student might ask whether sampling matters in an era of large administrative and behavioural datasets. It matters more. Large datasets are almost always found rather than designed — they represent whoever used a platform, filed a claim, or answered a call — and they carry coverage and self-selection properties that no amount of size corrects. The well-documented failures of large non-probability samples relative to small probability samples are now standard teaching cases precisely because the intuition runs the other way. The person in the room who can explain why the frame is the problem, and what a defensible design would have cost, is doing something the data volume cannot do for anyone.

AI Integration

This is a course whose central lesson is directly relevant to how artificial intelligence systems are built and evaluated, which makes the AI material substantive rather than decorative.

The training data of a machine learning system is a sample, and almost never a probability sample. Every argument the course makes about coverage error, self-selection and frame defects applies to it without modification. A model trained on internet text is trained on the population of people who write on the internet, weighted by how much they write — a sampling frame with severe and well-documented coverage properties. A clinical model trained on one hospital system's records is trained on that system's catchment population. The course gives students the vocabulary to say precisely what is wrong with these, which is more useful than the general observation that models can be biased. "What is the target population, what is the frame, and what is the difference between them" is the right first question about any model, and this is the course that teaches you to ask it.

Synthetic respondents are the live controversy in this field and belong in the course. There is an active research programme using large language models to simulate survey respondents — generating answers conditioned on demographic profiles and treating the output as a substitute for, or supplement to, collected data. The appeal is obvious given falling response rates and rising field costs. The problems are the ones this course is built to identify: the model reproduces the central tendencies of how demographic groups are written about in its training data, it compresses within-group variance dramatically, it cannot represent anyone whose views are underrepresented in text, and it has no mechanism by which a genuinely new opinion could appear. Most seriously, design-based inference is licensed by the randomisation, and there is no randomisation here — the confidence interval around a synthetic estimate has no defensible interpretation. This is a good subject for a course paper, and students should be able to argue it in both directions.

For the working parts of the course, language models are useful with supervision. They are effective at explaining a derivation you are stuck on, at writing and debugging R survey or SAS SURVEY procedure code, at drafting questionnaire items for you to critique, and at summarising a long methodology report. They are unreliable on the details that matter most in this subject: they routinely confuse the variance formulas for stratified and cluster designs, omit the finite population correction, apply Neyman allocation where proportional allocation was specified, and — most damagingly — produce analysis code that ignores the design entirely, computing a standard error as though a complex sample were a simple random sample. That error does not throw an exception. It produces a plausible number with a standard error that is too small, and a confidence interval that is too narrow, and nothing about the output announces the problem.

The check is the one the course teaches, and it is worth stating as a rule: compute the design effect and confirm the software used the design. If a cluster sample's reported standard error is not larger than the simple-random-sample standard error, the design was ignored — whoever or whatever wrote the code. Verify against a worked example with known answers before trusting generated analysis code on real data. Follow your instructor's syllabus on permitted use.


Generated September 5, 2026 · Updated September 5, 2026