Author Archives: gronke

UOCAVA Report and data released

The Election Assistance Commission’s UOCAVA report and data have been released and are available here: http://www.eac.gov/research/uocava_studies.aspx

The data are released in DBF (three files), XLS, and Spss “sav” file formats.  My first glance indicates that it should be straightforward to link the files to the NVRA data the FIPSCODE variable.

The headline of the EAC press release trumpets the 93.2% success rates for ballots–meaning that percent of ballots cast were counted.

That’s the good news.  The less good news is that the number of ballots transmitted dropped nearly 40% from 2008.  By comparison, turnout among the general public, according to figures compiled by Michael McDonald, dropped 34%.

Use of the Federal Write In Absentee Ballot (FWAB) remains tiny–only 2% of total UOCAVA ballots (4294) and of all FWABs, 16% are being rejected.  Something is up, either with the format of the FWAB, the instructions, or something else that is causing voter error.  (Unfortunately, the questionnaire does not break down the reasons for reject by regular absentee vs. FWAB).

Finally, response rates in some areas are down slightly from the high rates in 2008, while response rates in other areas are up (see pg. 4 of the report).  I generated that table in 2008, and I’m glad to see they repeated it since getting response rates at 90% or above are critical to a high quality data collection effort.

Colorado Absentee Ballot Fight: Data Can Help This!

In the ongoing battle over absentee ballots in Colorado, we’ve heard the claims about disenfranchised military voters and we’ve heard the charges about partisanship.

Unfortunately, what we haven’t heard is some hard factual information that compares ballot return rates among active and inactive voters. Andrew Cole, spokesperson for Secretary of State Scott Gessler is quoted as saying “there were thousands of ballots mailed out to inactive voters in 2010 that were unaccounted for.”

I’ve tried to answer this question at the Denver County elections office. Total registration, active and inactive, was 297,558 according to the spreadsheet available here:

Of that total, 22,696 are “Inactive – Fail to vote”, or 7.63% of the total.

The number of mail ballots issued was 160,363, of which 121, 538 were returned and verified. 128,997 mail ballots were returned in total, leaving 31,366 total outstanding unreturned mail ballots, or 19.55% of the total.

What is unknown is whether this number is high or low, and whether the proportion is higher or lower among active vs. inactive voters. If we assume the proportions apply across the groups, then there would have been approximately 2400 mail ballots “unaccounted for” that were sent to inactive voters, with the remainder (nearly 29,000) sent to active voters.

While speculative—it is likely that the proportion of unreturned ballots is higher among inactive voters—these figures speak directly to the claim being made by Andrew Cole. And the problem of unaccounted for ballots, if viewed this way, is obviously much greater among active voters

Data visualization toolkit

This is a cross posting from the Statistical Modeling, Causal Inference, and Social Science blog. I’m trying to figure out whether some of these tools may help us scrape and aggregate elections data. The Clark County voter history file comes up in a search, for instance, but there are no dataset associated with the link.

How to move from Google spreadsheet to R (a widely used, open source statistical package)

Google refine, a tool for cleaning messy data.

Infochimps.  I have no idea what this is for.

Google Fusion Tables (I tried to load the NVRA data here, and it worked ok.. sort of .. county names didn’t match up very well).

I’m pretty sure it won’t be long before someone is going to start creating feeds from election returns into one or more of these sources.

Mapping the NVRA data to examine data coverage

I have it on good authority that Doug Chapin is going to highlight one way to isolate data anomalies.

Another way is to put these data on a map. Unlike a box plot, the map doesn’t let us easily see how much variation there is across a state, although the color map does provide some indication. Nor are values above 1.0 (which should be out of bounds) easily identified in this map.

However, the map does provide a very nice overview, and each county’s values can be identified by floating the mouse pointer over a county. In the map below, I plot the percent of registration forms that came from the DMV (QA6d) expressed as a proportion of all valid forms received (QA5a).

Light blue identifies states that reported either no data (South Dakota) or a single statewide value (e.g. Alaska).

The white spots are the areas of concern because these are counties where no data at all have been reported. Highly variable color maps (e.g. New York) indicate that some counties are doing a very good job identifying and reporting registration forms submitted by the DMV while other counties are lagging.

The map can be viewed here. Screen shot 2011-09-14 at 1.14.05 PMIgnore the reference to “unemployment” in the title bar.

Posting additional data should be relatively straightforward; I invite comments below.

An easier solution to the leading zero problem

  1. Start with the Excel file
  2. Add in 3 new columns (D, E, and F)
  3. Highlight column C and select “data / text to columns” from the menu
  4. Select “fixed format” and then use the cursor to separate columns 1-2, 3-5, and 6-9
  5. Finish

The result should look like this and allows a match/merge by State, County, and Township.

State Jurisdiction FIPSState FIPSCounty FIPSextra QA1a
AK ALASKA 02 0 0 560146
AL AUTAUGA COUNTY 01 1 0 34727
AL BALDWIN COUNTY 01 3 0 114952
AL BARBOUR COUNTY 01 5 0 16450
AL BIBB COUNTY 01 7 0 12239
AL BLOUNT COUNTY 01 9 0 31874
AL BULLOCK COUNTY 01 11 0 7650
AL BUTLER COUNTY 01 13 0 12898

Leading zero problems in the NVRA Dataset

In the EAC’s NVRA 2010 dataset, the FIPS code is stored as a numeric value, and this will cause problems with anyone who is trying to match / merge this with other data because 10 state FIPS codes start with a “0”.  Most statistical programs will strip that leading zero, making it hard to identify the states properly.

For instance, lines 2-4 of the Excel file look like this:

AK ALASKA 200000000 560146 only active voters
AL AUTAUGA COUNTY 0100100000 34727 active and inactive registered voters
AL BALDWIN COUNTY 0100300000 114952 active and inactive registered voters

The third column contains the FIPS code; columns 1-2 contain the state code, columns 3-5 the county code, an additional columns refer to smaller jurisdictions (mainly townships in NE and the Midwest).

The first problem you’ll notice is that Alaska’s FIPS is already artificially shortened to “2” not “02” an users will need to fix this manually.  Next, additional FIPS codes will come across a “10010000” (for example) without the leading zero.

The solution for now is a kludge; insert a dummy line into the Excel file and include a string value such as an “x” in column 3.  This will trick Excel into outputting the values as text, and statistical programs will input the data as character strings.  Then the process is straightforward; capture substrings (01, 02, etc) and match / merge away!