THE CDS-2014 DATA FILE STRUCTURE

From PSIDWiki
Jump to navigation Jump to search

CDS-2014 data are organized into five individual study components and two supporting files:

1. Primary Caregiver Household Interview (one record per primary caregiver) 2. Primary Caregiver Child Interview (one record per CDS-2014 child) 3. Child Interview (one record per age-eligible CDS-2014 child) 4. Child Assessments (one record per age-eligible CDS-2014 child in families selected for the in-home component) 5. Time Diary (one record per CDS-2014 child in families selected for the in-home component), organized into three files: a. Time Diary Activity File (one record per activity) b. Time Diary Aggregated Activity File (one record per child) c. Time Diary Questionnaire Administration File (one record per child) 6. Demographic File 7. ID Map File

CDS-2014 data are contained in nine files—one file for each individual study component or supporting file except for the time diary data which are organized into three files. Table 2.1 summarizes the five study components and two supporting files according to the CDS-2014 individual for whom the data are available and also lists the number of records in each component/file.

Primary Caregiver and Child-Level Data Files

The PCG-Household file is released at the caregiver level. This means that each caregiver is represented by one record on the PCG-Household file. There are 2,517 records in this file, with one record corresponding to each PCG in the CDS-2014 sample. A PCG may be a caregiver to more than one CDS-2014 child, so information from one record on the PCG-Household file can be connected to multiple child records on a child-level file.

All other files are released at the child level. This means that each CDS-2014 child who participated in a given study component is represented by one record on the corresponding data file. (The time diary activity file is an exception.) Because of this design, a data extract from the PSID Online Data Center that includes variables from both the PCG-Household Interview and variables from any child-level file will include two content data files: one PCG-Household data file and one child-level file.

Unique Person Identifiers in the PCG-Household and Child-Level Files

The unique identifiers on the PCG-Household File identify the primary caregiver and the unique identifiers that appear on the child-level file refer to the child. Specifically, the 1968 ID (ER30001), person number (ER30002), 2013 family interview identifier (ER34201), 2013 family interview sequence number (ER34202), and 2013 relationship to head (ER34203) all refer to the primary caregiver on the PCG-Household file and to the child on the child-level file.

A mapping file is provided with all CDS-2014 data extracts in order to link data from the PCG- Household file to data from any child file. The ID Map File enables a one-to-many merge between the PCG-Household file and any child-level file. It contains 4,353 records, one for each child for whom information is available in CDS-2014. The following variables are included in the ID Map File:

ER30001 – Child 1968 interview number ER30002 – Child’s person number PCGID68 – 1968 interview number for the child’s 2014 primary caregiver PCGPN – Person number for the child’s 2014 primary caregiver CDSHID – CDS-2014 household interview number CHLDINST14 – Child’s sequence number (roster position) in CDS-2014 PCGINST14 – Primary caregiver’s sequence number in CDS-2014 Identifiers to Use for Merging PCG Data with Child Data

In the ID Map File, the variables ER30001 and ER30002 together provide unique values for every child on the mapping file. The variables PCGID68 and PCGPN provide a value for each PCG who has been observed previously in PSID and has an assigned value 1968 ID and person number (see below for an important note).

To allow users to merge information between the PCG-HH file and any child-level file, the ID Map File includes a CDS household interview number (CDSHID). This variable is similar to the family interview ID number assigned to a family unit in a given wave of the Core PSID interview. It uniquely distinguishes individual PCG-Household interviews in CDS-2014, but it does not correspond to any information outside of CDS-2014.

Merging PCG data with Child Data

To merge data between the PCG-Household File and any child-level file using the unique CDS household interview number, users may take the following steps:

1. Conduct a one-to-one merge between the child-level file and the ID Map File using ER30001 and ER30002 as the unique identifiers. This merge will add the household interview ID and PCG identifiers to the child-level file. 2. In the PCG-Household File, rename ER30001 and ER30002 to PCGID68 and PCGPN respectively. Users may also wish to drop or rename the variables ER34201, ER34202, and ER34203, which uniquely identify the caregiver in the 2013 Core PSID data. This will prevent values on these variables from being overwritten (or overwriting the values of these unique identifiers for children) when the file is merged to the child-level file in the next step. In addition, rename H14CDSHID to CDSHID and rename H14INST to PCGINST14. These will be the merging variables in the next step, so the variable names will need to be the same in this file and the ID Map File. 3. Conduct a one-to-many merge between the PCG-Household File and the ID Map File using CDSHID and PCGINST14 as the unique identifiers. This will put the PCG- Household data, PCG identifiers, and household interview identifier at the child level. 4. Conduct a one-to-one merge between (a) the new child-level file that contains PCG- Household data created in Step 3 and (b) the child-level file containing child data. Use ER30001 and ER30002 as the unique identifiers. Be aware that in some cases child data were collected but no PCG-Household interview was completed. For these cases, ER30001 and ER30002 will not take a value on the new child-level PCG-Household file created in Step 3. Users may wish to resolve this by deleting these records from the child-level PCG-Household file or by conducting a many-to-one merge at this step.

The preceding steps will produce a child-level data set that includes PCG-Household information for all cases where those data are available. In addition, as noted above, it may include records from households where only a PCG-Household interview was completed and no child data were collected or where child data were collected but no PCG-Household data were collected.

IMPORTANT NOTE. In 65 cases in CDS-2014, the primary caregiver was not observed in the child’s household at the time of the 2013 Core PSID interview (and had never been observed in PSID previously). As a result, the primary caregiver has no assigned 1968 ID or person number. Users should be aware that no other data from the PSID Online Data Center may be attached to the 65 cases where a primary caregivers lacks a unique identifier.