|
|
| Line 15: |
Line 15: |
| For analysts using the Child Interview data, special care must be taken to choose the correct weight, because Child Interviews for one segment of the sample (children aged 8–11 years) were only conducted as part of the in-home component. Table 5.1 describes the appropriate weight to use based on the ages of children with Child Interview data being analyzed. Note that for analyses based on all children aged 8–17 who completed the Child Interview, the analyst must create a new variable with values equal to X14IHWGT for children aged 8–11 years and X14CHWGT for children aged 12–17 years. | | For analysts using the Child Interview data, special care must be taken to choose the correct weight, because Child Interviews for one segment of the sample (children aged 8–11 years) were only conducted as part of the in-home component. Table 5.1 describes the appropriate weight to use based on the ages of children with Child Interview data being analyzed. Note that for analyses based on all children aged 8–17 who completed the Child Interview, the analyst must create a new variable with values equal to X14IHWGT for children aged 8–11 years and X14CHWGT for children aged 12–17 years. |
| We provide further guidance on how researchers can conduct weighted analyses of other CDS- 2014 components at the end of this chapter. | | We provide further guidance on how researchers can conduct weighted analyses of other CDS- 2014 components at the end of this chapter. |
|
| |
| Overview of Method to Construct CDS-2014 Child Weights
| |
|
| |
| We first constructed the CDS-2014 Main Child Weight, which was then modified to produce the CDS-2014 Child In-Home Weight. The basic steps to producing these weights were as follows:
| |
|
| |
| 1. Account for all probabilities of selection for eligible families and children through to the initial determination of eligibility for CDS-2014.
| |
| 2. Adjust for CDS-2014 non-response.
| |
| 3. Set aside the small number of CDS-2014 cases residing outside the U.S., for which their CDS-2014 Main Child Weight is now final.
| |
| 4. Post-stratify the attrition-adjusted CDS-2014 sample selection weights to 2014 American Community Survey (ACS) adjusted population totals based on year of birth, gender, race, and Census region.
| |
| 5. Trim very large and very small values of the post-stratified weights.
| |
| 6. Post-stratify the trimmed weights to produce the final CDS-2014 Main Child Weight for the cases currently residing in the U.S.
| |
| 7. Pool the U.S cases and non-U.S. cases for CDS-2014, which both now have their final Main Child Weight.
| |
| 8. Adjust the final CDS-2014 Main Child Weight to produce the CDS-2014 Child In-Home Weight.
| |
|
| |
| Method to Construct CDS-2014 Child Weights
| |
| We next describe the steps of the process for constructing the CDS-2014 weights. Step 1. Selection Probabilities for CDS-2014
| |
| For all eligible CDS-2014 child cases (n=5,816), a base probability of selection weight was established using the household weight from the 2013 Core PSID.
| |
|
| |
| Step 2. Non-Response Adjustment
| |
|
| |
| A non-response adjustment factor for the weight was obtained from a logistic regression model of the response outcome. All eligible CDS-2014 child cases were included in the model. Data from the 2013 Core PSID were used as covariates in the model predicting an indicator of non- response, y, with y=0 if the case was non-response and y=1 if the case was coded as complete. The estimated coefficients and standard errors for the logistic regression model are reported in Table 5.3.
| |
|
| |
| The results indicate that the probability of response in CDS-2014 was higher among African Americans and whites, in households headed by women, and in households with fewer children; the probability of response was lower in households with missing information about the head’s education and in households outside the U.S. Although a number of variables in the response propensity model that are not statistically significant predictors of CDS-2014 response (e.g., household income, metro status, and Census region), these non-significant variables were retained in the model used to derive estimates of the propensity of response. Overall, the Hosmer-Lemeshow test of goodness of fit test (X2=11.47, 8 df, p=0.18) suggests that the response model provides an acceptable fit.
| |
|
| |
| Based on the estimated logistic regression model, predicted probabilities of response were computed for each CDS sample case and grouped into deciles. These decile groups served as the classes within which a uniform non-response weighting adjustment was applied.13 Each respondent case was assigned a non-response adjustment factor equal to the inverse of the median predicted probability of successful CDS 2014-interview within its decile weighting class. The median response propensity and adjustment factor for each decile of the predicted probability response in CDS-2014 are shown in Table 5.4.
| |
|
| |
| The probability of selection weight for each CDS-2014 observation was then multiplied by the non-response adjustment factor to produce an interim weight that adjusts for probability of selection and CDS-2014 nonresponse.
| |
|
| |
| Step 3. Non-U.S. Cases
| |
|
| |
| There were 30 eligible cases in CDS-2014 that resided outside the U.S. during the fieldwork period. Although interviews were attempted for all of these cases and completed among some of them, these cases are not included in the post-stratification because the control total for the post-stratification process are based on the U.S. resident population. At this step, for the non-U.S. CDS-2014 cases, the Main Child Weight is designated to be complete.
| |
|
| |
| Step 4. Post-Stratification to Population Control Totals
| |
|
| |
| We next post-stratified the CDS-2014 interim weights from Step 2 to population control totals
| |
|
| |
| calculated using data from the 2013 American Community Survey. Strata were formed based on the following respondent characteristics:
| |
|
| |
| • Child sex (male/female)
| |
| • Birth year of child (1997–2013)
| |
| • Child race/ethnicity (Hispanic, non-Hispanic black, non-Hispanic white or other)
| |
| • Census region (Northeast, Midwest, South, West)
| |
|
| |
| Strata defined by the full four-way cross-classification of these categorical variables were collapsed as needed to ensure a minimum count of approximately 15–20 individuals in each cell. Table 5.5 shows the CDS sample count and CDS weighted estimates, the ACS population estimates, and the post-stratification adjustment factors for each of the 105 cells defined by birth year, sex, race/ethnicity, and region. Note that CDS-2014 excluded some children born in early 1997 (who were selected to participate in the original CDS) and some children born in late 2013 (after the 2013 Core PSID interview was completed). The ACS control totals shown in Table
| |
| 4.5 and used in the post-stratification weighting have been constructed to exclude children born outside the CDS-2014 eligibility window in 1997 and 2013.
| |
|
| |
| The initial post-stratification adjustment factors were computed as the ratio of the ACS control totals to the CDS-2014 weighted population estimate count (using the interim weight from Step 2). The initial post-stratification adjustment factors were then applied to the interim weight to produce an initial post-stratified weight.
| |
|
| |
| Step 5. Trimming of Weights
| |
|
| |
| The distribution of the interim, post-stratified weights was examined and a decision was made to trim extreme values at each end of the distribution. The reason for trimming the weights is to reduce the influence of extreme weight values on the variances of sample estimates of population statistics. Trimming the weight distribution also provides some protection against arbitrary combinations of extreme weights and large or unique values of substantive variables that could exert high leverage on multivariate analyses such as regression modeling. The trimming rule applied to the CDS-2014 Main Child Weight assigned the cases with the weight values in the top two percent and in the bottom two percent of the weight distribution to the 98th and 2nd percentile values of the weight distribution, respectively.
| |
|
| |
| Step 6. Post-Stratification after Trimming of Weights
| |
|
| |
| After trimming the weights, the post-stratification procedure (Step 4) was repeated so that the final trimmed weights again matched the desired ACS population control totals.
| |
|
| |
| Step 7. Combining the U.S. and Non-U.S. Cases
| |
|
| |
| The final step in creating the Main Child Weight is to combine the weights from Step 6 for cases in the U.S with the weights from Step 3 for the non-U.S. cases.
| |
|
| |
| Step 8. Produce the CDS-2014 In-Home Child Weight
| |
|
| |
| Half of the CDS-2014 families were randomly selected to receive an in-home visit. The selection process was probability-based, but all cases did not have an equal chance of selection for the in-home survey administration. To account for the subsampling of CDS families for in- home interview administration, the CDS-2014 Child In-Home Weight includes an additional sample selection adjustment to the Main Child Weight. These adjusted weights were then post- stratified to the ACS population control totals using a process identical to that described in Steps 4–6 above, except that due the smaller sample size for the in-home interviews a further collapsing of strata was necessary. Census region was dropped entirely from the post- stratification scheme and wider birth-year cells were used. The revised post-stratification scheme is shown in Table 5.6.
| |
|
| |
| Method to Construct CDS-2014 PCG Weights
| |
|
| |
| The CDS-2014 PCG weight was derived entirely from the CDS-2014 Main Child Weight. In particular, PCGs in CDS-2014 were assigned the average weight over all children for whom they were the responsible primary caregiver. For PCGs for whom there were no corresponding children in the sample (because no child interview components were completed and hence no child weight was constructed), a PCG weight was calculated based on imputed values for the missing child weights. Missing child weights were imputed based on predicted values from a regression model that included covariates from the child-level non-response model described above and the 2013 Core PSID weight.
| |
|
| |
| Summary of CDS-2014 Weights
| |
|
| |
| In Table 5.7 we list the CDS-2014 child weights and the PCG weight and present case counts and summary statistics. that the case count for the Main Child Weight (4,333) is approximately twice the count for the Child In-Home Weight, reflecting the fact that approximately half of the CDS-2014 sample was selected for the in-home component. The mean weight for the Child In- Home Weight (28,241.22) is correspondingly about twice as large as the mean weight for the Main Child Weight (14,061.74), reflecting the fact that these two samples both weight to the same national population of children. Note also that the weighted total population from both CDS-2014 samples is approximately 61 million children, which is about 83 percent of the estimated U.S. population of children aged 0–17 years of 73 million in 2013. The PSID count is lower primarily because it excludes children in post-1997 immigrant families and includes only about half of children in the youngest and oldest single-year age groups.
| |
|
| |
| The case count for the PCG weight is 2,517 and the mean weight is 14,018.62, which (by construction) is very similar to the mean for Main Child Weight (14,061.74). The weighted total population of PCGs in CDS-2014 is 35 million.
| |
|
| |
| Recommendations for Using the CDS-2014 Weights
| |
|
| |
| In this section, we summarize our recommendations for using the CDS-2014 weights. Our basic recommendation is for data users to always use the provided weights in all of their analyses. In addition, we recommend that, when calculating standard errors, data users should wherever possible account for the clustering of the CDS-2014 data by family. Standard errors should reflect the fact that siblings are more likely to have similar outcomes and characteristics than children selected at random. Controlling for family-level clustering of siblings also provides an appropriate correction due to clustering of families by household or neighborhood and recognizes the fact that generally it is only possible to control for a single level of clustering. Furthermore, when analyses focus on a subset of children (from either the full sample or the in- home component), data users should use an appropriate “sub-population” adjustment. Clustering-corrected standard errors and sub-population commands are available in most standard statistical software (including SAS and Stata).
| |
|
| |
| Main Child Weight (X14CHW GT). This weight should be used for all analyses in which the full sample of children in CDS-2014 are the focus of the analysis. This is the weight to use with data from the PCG Child Instrument or for data on children aged 12–17 years from the Child Instrument.
| |
|
| |
| Child In-Home Weight (X14IHWGT). This weight should be used for all analyses in which the analysis focuses on measures available only in the in-home component of CDS-2014, which includes the following components: Child Interviews for children aged 8–11 years, the Woodcock-Johnson Tests of Achievement in reading and math, and the CDS Time Diaries.
| |
|
| |
| PCG Weight (H14PCGWGT). This weight should be used for all analyses in which the sample of PCGs in CDS-2014 are the focus of the analysis. This is the weight to use with data from the PCG Household Instrument or for other data on PCGs.
| |
|
| |
| Finally, if users have questions about whether their analyses should be weighted or unweighted or about how to reflect the sampling design in their calculation of parameter estimates and standard errors, they should consult with a survey statistician.
| |