This assignment explores web data collection using rvest, from a simple scraping sandbox to a Wikipedia table and a comparison with a structured World Bank data source. The final section applies these methods to a research plan on teacher shortages in North Texas.
NoteWhat this assignment covers
Checking permission before making a web request
Scraping and combining data from multiple pages
Cleaning a real-world HTML table
Comparing scraped data with an API
Designing a responsible web data collection plan
0. Permission Check
Before scraping Wikipedia, I checked whether the page could be accessed according to its robots.txt rules using robotstxt::paths_allowed().
library(rvest)library(dplyr)
Attaching package: 'dplyr'
The following objects are masked from 'package:stats':
filter, lag
The following objects are masked from 'package:base':
intersect, setdiff, setequal, union
library(stringr)library(readr)
Attaching package: 'readr'
The following object is masked from 'package:rvest':
guess_encoding
library(tidyr)library(polite)library(janitor)
Attaching package: 'janitor'
The following objects are masked from 'package:stats':
chisq.test, fisher.test
The requested path was permitted, so I proceeded with the Wikipedia request.
1. Warm-Up: Books to Scrape
I first practiced the scraping workflow on Books to Scrape, a sandbox website specifically designed for practicing web scraping. Because the site is static and contains no JavaScript requirement, it provides a low-risk environment for learning the basic rvest functions.
# A tibble: 6 × 4
title price rating price_gbp
<chr> <chr> <chr> <dbl>
1 A Light in the Attic £51.77 Three 51.8
2 Tipping the Velvet £53.74 One 53.7
3 Soumission £50.10 One 50.1
4 Sharp Objects £47.82 Four 47.8
5 Sapiens: A Brief History of Humankind £54.23 Five 54.2
6 The Requiem Red £22.65 One 22.7
This extracts the book title, price, and rating from each product card. The price is also converted from text to a numeric variable for later analysis.
1b. Scraping Five Pages
After testing the extraction on one page, I turned it into a function so that the same process could be applied to five pages.
I next applied the same rvest workflow to Wikipedia’s table of foreign-exchange reserves. Unlike the Books to Scrape sandbox, this was a real-world table with a more complicated structure.
wiki_url <-paste0("https://en.wikipedia.org/wiki/","List_of_countries_by_foreign-exchange_reserves")wiki_pg <-read_html(wiki_url)## Never assume the table you want is [[1]]. Look first.tabs <- wiki_pg |>html_elements("table.wikitable")length(tabs) # how many candidates? --> 2
# A tibble: 6 × 4
country_as_recognized_by_…¹ continent foreign_exchange_res…² last_reporteddate
<chr> <chr> <dbl> <chr>
1 China Asia 3854885 1 August 2026
2 Japan Asia 1287470 30 June 2026
3 Switzerland Europe 939140 30 June 2026
4 India Asia 785710 04 September 2026
5 Russia Europe/A… 761200 21 August 2026
6 Taiwan Asia 597150 30 June 2026
# ℹ abbreviated names: ¹country_as_recognized_by_the_un,
# ²foreign_exchange_reserves
2b. A Second Wikipedia Table
To test whether the same code would work on a different table, I used the FIFA World Cup results table from Wikipedia.
# A tibble: 10 × 8
`World Cup` `Golden Ball` `Golden Boot` Goals `Golden Glove` `Clean sheets`
<chr> <chr> <chr> <int> <chr> <chr>
1 1930 Uruguay Not Awarded Guillermo St… 8 Not Awarded N/A
2 1934 Italy Not Awarded Oldřich Neje… 5 Not Awarded N/A
3 1938 France Not Awarded Leônidas 7 Not Awarded N/A
4 1950 Brazil Not Awarded Ademir 9 Not Awarded N/A
5 1954 Switzer… Not Awarded Sándor Kocsis 11 Not Awarded N/A
6 1958 Sweden Not Awarded Just Fontaine 13 Not Awarded N/A
7 1962 Chile Not Awarded Flórián Albe… 4 Not Awarded N/A
8 1966 England Not Awarded Eusébio 9 Not Awarded N/A
9 1970 Mexico Not Awarded Gerd Müller 10 Not Awarded N/A
10 1974 West Ge… Not Awarded Grzegorz Lato 7 Not Awarded N/A
# ℹ 2 more variables: `FIFA Young Player Award` <chr>,
# `FIFA Fair Play Trophy` <chr>
The scraping step itself still worked, but the cleaning code from the foreign-exchange table did not transfer directly. The FX table contained columns such as country_as_recognized_by_the_un and a two-row header, while the FIFA table had columns such as world_cup, golden_ball, and golden_boot. Therefore, cleaning commands that referenced the FX-specific columns could not be reused unchanged.
WarningWhat broke?
The extraction worked, but the cleaning code was tied to the structure of the original table. The FIFA table had different column names and did not contain the same multi-row header or Total/World rows. This showed that scraping code is often dependent on the specific structure of a webpage.
# A tibble: 6 × 9
world_cup golden_ball golden_boot goals golden_glove clean_sheets
<chr> <chr> <chr> <int> <chr> <chr>
1 1930 Uruguay Not Awarded Guillermo Stábile 8 Not Awarded N/A
2 1934 Italy Not Awarded Oldřich Nejedlý 5 Not Awarded N/A
3 1938 France Not Awarded Leônidas 7 Not Awarded N/A
4 1950 Brazil Not Awarded Ademir 9 Not Awarded N/A
5 1954 Switzerland Not Awarded Sándor Kocsis 11 Not Awarded N/A
6 1958 Sweden Not Awarded Just Fontaine 13 Not Awarded N/A
# ℹ 3 more variables: fifa_young_player_award <chr>,
# fifa_fair_play_trophy <chr>, year <dbl>
3. Now Do It the Right Way
Wikipedia was useful for practicing scraping, but the same information is available through a structured World Bank indicator. This provides an opportunity to compare a scraped source with an API-based source.
country iso2c iso3c year FI.RES.TOTL.CD status lastupdated
1 Afghanistan AF AFG 2025 NA 2026-07-13
2 Afghanistan AF AFG 2024 NA 2026-07-13
3 Afghanistan AF AFG 2023 NA 2026-07-13
4 Afghanistan AF AFG 2022 NA 2026-07-13
5 Afghanistan AF AFG 2021 NA 2026-07-13
6 Afghanistan AF AFG 2020 9748946327 2026-07-13
region capital longitude latitude
1 Middle East, North Africa, Afghanistan & Pakistan Kabul 69.1761 34.5228
2 Middle East, North Africa, Afghanistan & Pakistan Kabul 69.1761 34.5228
3 Middle East, North Africa, Afghanistan & Pakistan Kabul 69.1761 34.5228
4 Middle East, North Africa, Afghanistan & Pakistan Kabul 69.1761 34.5228
5 Middle East, North Africa, Afghanistan & Pakistan Kabul 69.1761 34.5228
6 Middle East, North Africa, Afghanistan & Pakistan Kabul 69.1761 34.5228
income lending
1 Low income IDA
2 Low income IDA
3 Low income IDA
4 Low income IDA
5 Low income IDA
6 Low income IDA
FI.RES.TOTL.CD measures total reserves, including gold, in current U.S. dollars.
3b. Comparison
library(lubridate)
Attaching package: 'lubridate'
The following objects are masked from 'package:base':
date, intersect, setdiff, union
# Convert Wikipedia's reporting date from character to Datereserves_compare <- reserves |>mutate(last_reporteddate =dmy(last_reporteddate),# Wikipedia reports reserves in US$ million,# so convert to current US dollarswiki_reserves_usd = foreign_exchange_reserves *1000000,# Extract the year from the reporting datewiki_year =year(last_reporteddate) ) |>select(country = country_as_recognized_by_the_un, last_reporteddate, wiki_year, wiki_reserves_usd )
Warning: There was 1 warning in `mutate()`.
ℹ In argument: `last_reporteddate = dmy(last_reporteddate)`.
Caused by warning:
! 1 failed to parse.
head(reserves_compare)
# A tibble: 6 × 4
country last_reporteddate wiki_year wiki_reserves_usd
<chr> <date> <dbl> <dbl>
1 China 2026-08-01 2026 3854885000000
2 Japan 2026-06-30 2026 1287470000000
3 Switzerland 2026-06-30 2026 939140000000
4 India 2026-09-04 2026 785710000000
5 Russia 2026-08-21 2026 761200000000
6 Taiwan 2026-06-30 2026 597150000000
## Match each country to the World Bank value for the same year ---------------reserves_comparison <- reserves_compare |>left_join( reserves_api |>select( country, year,world_bank_reserves_usd = FI.RES.TOTL.CD,world_bank_last_updated = lastupdated ),by =c("country"="country", "wiki_year"="year") ) |>mutate(difference_usd = wiki_reserves_usd - world_bank_reserves_usd,difference_percent = (difference_usd / world_bank_reserves_usd) *100 ) |>arrange(desc(abs(difference_usd)))
The two sources do not always report the same values for the same country and calendar year. The differences appear particularly large for some countries, showing that matching only by country and year does not necessarily mean that the two sources represent the same reporting vintage.
For example, Wikipedia reports U.S. reserves of approximately $253.8 billion with a reporting date of August 22, 2025, while the World Bank dataset reports approximately $1.39 trillion for 2025 and was updated in July 2026. This illustrates why the reporting date and dataset vintage matter when comparing values from different sources.
country year FI.RES.TOTL.CD lastupdated
1 United States 2025 1.385447e+12 2026-07-13
2 United States 2024 9.100365e+11 2026-07-13
3 United States 2023 7.734262e+11 2026-07-13
4 United States 2022 7.066442e+11 2026-07-13
5 United States 2021 7.161523e+11 2026-07-13
6 United States 2020 6.283697e+11 2026-07-13
7 United States 2019 5.167006e+11 2026-07-13
8 United States 2018 4.499071e+11 2026-07-13
9 United States 2017 4.512853e+11 2026-07-13
10 United States 2016 4.059424e+11 2026-07-13
11 United States 2015 3.837285e+11 2026-07-13
3c. When is scraping the right tool?
Scraping is the right tool when the information you need is publicly available online but there is no suitable API, downloadable dataset, or other structured way to access it. It should be used when it is the only practical way to obtain the data, while respecting the website’s robots.txt, terms of service, and limits on requests.
NoteMain lesson
- Wikipedia: edited by volunteers, sourced from many places, no schema,
no versioning, no uncertainty, changes without notice
- World Bank: one compiler, documented methodology, stable indicator code
Scraping was the wrong tool here. Knowing when it is the only tool is the
skill this course is actually teaching.
4. Data Plan: Measuring Teacher Shortages in Dallas County ISDs
Research Question
Do administrative workforce indicators and online teacher vacancy records identify similar staffing challenges across public school districts in Dallas County?
This project will compare standardized administrative workforce indicators from the Texas Education Agency (TEA) with independently collected teacher vacancy records from district recruitment websites. The goal is to examine whether different approaches to measuring teacher shortages produce similar assessments of district staffing conditions.
4a. Data Sources
Primary unit of analysis: Independent school district (ISD)
Raw web observation: Individual job posting
Study area: Dallas County
Source
What I would collect
Why use it?
Texas Education Agency
Teacher attrition, retention, employment, certification, enrollment, and other workforce indicators
Standardized administrative data across Texas districts
District recruitment websites
Job title, subject/position, posting date, district
Direct measure of positions districts are actively advertising
NCES
Contextual information such as teacher workforce and compensation
Additional context if needed
Texas Schools Project
Additional education data if needed
Potential supplemental source
TEA provides standardized administrative measures of the teacher workforce, making it useful for comparing districts. District job postings measure a different aspect of staffing conditions: positions that districts are actively seeking to fill. Comparing these sources allows the project to examine whether administrative indicators and vacancy data tell the same story.
TipWhy use two measures?
Administrative data: What happened to the teacher workforce?
Vacancy data: What positions are districts currently trying to fill?
The comparison is the main research contribution of the data collection strategy.
4b. robots.txt and Terms of Service
We will check each recruitment site’s robots.txt file before collecting data and will review the Terms of Service where available. We will only automate collection from publicly accessible pages when the relevant path is permitted, and we will not bypass access restrictions, authentication, or other technical protections.
Our initial checks show that several district sites permit access to their public job-listing pages, while Highland Park ISD’s recruitment site currently disallows scraping. Sites with unclear or timed-out robots.txt checks will be verified before collection. If automated scraping is not permitted, we will exclude the site from automated collection or use manual collection if allowed.
District
Recruitment platform
Site type
robots.txt / access
Irving ISD
SchoolSpring
Dynamic
Path allowed
Dallas ISD
Taleo
Dynamic
paths_allowed() returned TRUE
Plano ISD
Applitrack
Dynamic
Relevant paths appear permitted
Richardson ISD
Applitrack
Dynamic
Relevant paths appear permitted
Highland Park ISD
TX Ed Job Network
Dynamic
Scraping disallowed
Carrollton-Farmers Branch ISD
Applitrack
Dynamic
To be checked
Cedar Hill ISD
District/Skyward-style system
TBD
Requires further inspection
Coppell ISD
Winocular
Dynamic
Path allowed
DeSoto ISD
Teams
Dynamic
Path allowed
Duncanville ISD
Skyward
Dynamic
robots.txt retrieval worked; path check timed out
Garland ISD
Applitrack
Dynamic
robots.txt retrieval worked; path check had an error
Grand Prairie ISD
Skyward
Dynamic
robots.txt retrieval worked; path check timed out
Lancaster ISD
Applitrack
Dynamic
robots.txt retrieval worked; path check had an error
Mesquite ISD
Applitrack
Dynamic
To be checked
Sunnyvale ISD
Applitrack
Dynamic
URL/platform needs verification
Highland Park ISD is currently not suitable for automated scraping because the recruitment site’s robots.txt disallows scraping. I would exclude it from automated collection rather than attempt to bypass the restriction.
4c. Whether an API exists
TEA
TEA data will be collected through available administrative datasets or structured dashboard data rather than scraped from the website.
District vacancies
The district recruitment systems do not share one common API because districts use different vendors and platforms. Where a suitable API or downloadable dataset is unavailable, permitted public vacancy pages may be collected through automated or manual methods, such as scraping, depending on the site’s structure. However, more research is being done for the most ethical way of obtaining the sites’ information.
4d. Volume, rate limiting, and cache
The project will collect a relatively small number of district vacancy pages rather than continuously crawling entire websites. Requests will be rate-limited and spaced out, and previously collected pages and extracted records will be stored locally so that pages do not need to be repeatedly requested. Each observation will include a retrieval date to distinguish the collection date from the original job posting date.
I will initially test the method on a small subset of districts before scaling it to the full study area.
4e. What breaks when the page changes
Potential change
What could happen
How I would detect it
HTML structure changes
Scraper returns empty or incomplete data
Check number of records and expected columns
Recruitment platform changes
URL or selectors stop working
Test each district before collection
Job fields change
Posting date/title no longer extracted
Validate expected fields
Site becomes more dynamic
read_html() no longer captures listings
Compare rendered page with extracted HTML
Robots.txt changes
Automated collection may no longer be permitted
Re-check robots.txt before collection
Old postings disappear
Historical vacancy counts become incomplete
Record retrieval date and archive collected results
Unexpected changes in the number of records or missing expected fields would trigger manual inspection rather than allowing the scraper to silently produce an incomplete dataset.