Assignment 4: Webscraping 1

This assignment explores web data collection using rvest, from a simple scraping sandbox to a Wikipedia table and a comparison with a structured World Bank data source. The final section applies these methods to a research plan on teacher shortages in North Texas.

NoteWhat this assignment covers
  1. Checking permission before making a web request
  2. Scraping and combining data from multiple pages
  3. Cleaning a real-world HTML table
  4. Comparing scraped data with an API
  5. Designing a responsible web data collection plan

0. Permission Check

Before scraping Wikipedia, I checked whether the page could be accessed according to its robots.txt rules using robotstxt::paths_allowed().

library(rvest)
library(dplyr)

Attaching package: 'dplyr'
The following objects are masked from 'package:stats':

    filter, lag
The following objects are masked from 'package:base':

    intersect, setdiff, setequal, union
library(stringr)
library(readr)

Attaching package: 'readr'
The following object is masked from 'package:rvest':

    guess_encoding
library(tidyr)
library(polite)
library(janitor)

Attaching package: 'janitor'
The following objects are masked from 'package:stats':

    chisq.test, fisher.test
library(WDI)
robotstxt::paths_allowed(
  "https://en.wikipedia.org/wiki/List_of_countries_by_foreign-exchange_reserves")

 en.wikipedia.org                      
[1] TRUE

The requested path was permitted, so I proceeded with the Wikipedia request.

1. Warm-Up: Books to Scrape

I first practiced the scraping workflow on Books to Scrape, a sandbox website specifically designed for practicing web scraping. Because the site is static and contains no JavaScript requirement, it provides a low-risk environment for learning the basic rvest functions.

1a. Scraping One Page

books_pg <- read_html("https://books.toscrape.com/catalogue/page-1.html")

books <- tibble(
  title  = books_pg |> html_elements("article.product_pod h3 a") |> html_attr("title"),
  price  = books_pg |> html_elements("article.product_pod p.price_color") |> html_text2(),
  rating = books_pg |> html_elements("article.product_pod p.star-rating") |>
    html_attr("class") |> str_remove("star-rating ")
) |>
  mutate(price_gbp = parse_number(price))

head(books)
# A tibble: 6 × 4
  title                                 price  rating price_gbp
  <chr>                                 <chr>  <chr>      <dbl>
1 A Light in the Attic                  £51.77 Three       51.8
2 Tipping the Velvet                    £53.74 One         53.7
3 Soumission                            £50.10 One         50.1
4 Sharp Objects                         £47.82 Four        47.8
5 Sapiens: A Brief History of Humankind £54.23 Five        54.2
6 The Requiem Red                       £22.65 One         22.7

This extracts the book title, price, and rating from each product card. The price is also converted from text to a numeric variable for later analysis.

1b. Scraping Five Pages

After testing the extraction on one page, I turned it into a function so that the same process could be applied to five pages.

scrape_books_page <- function(page) {
  
  url <- paste0(
    "https://books.toscrape.com/catalogue/page-",
    page,
    ".html"
  )
  
  pg <- read_html(url)
  
  tibble(
    title = pg |>
      html_elements("article.product_pod h3 a") |>
      html_attr("title"),
    
    price = pg |>
      html_elements("article.product_pod p.price_color") |>
      html_text2(),
    
    rating = pg |>
      html_elements("article.product_pod p.star-rating") |>
      html_attr("class") |>
      str_remove("star-rating ")
  ) |>
    mutate(
      price_gbp = parse_number(price),
      page = page
    )
}
# 5 variables: title, price, rating, price_gbp, page

books_5pages <- lapply(1:5, scrape_books_page) |>
  bind_rows()

glimpse(books_5pages)
Rows: 100
Columns: 5
$ title     <chr> "A Light in the Attic", "Tipping the Velvet", "Soumission", …
$ price     <chr> "£51.77", "£53.74", "£50.10", "£47.82", "£54.23", "£22.65", …
$ rating    <chr> "Three", "One", "One", "Four", "Five", "One", "Four", "Three…
$ price_gbp <dbl> 51.77, 53.74, 50.10, 47.82, 54.23, 22.65, 33.34, 17.93, 22.6…
$ page      <int> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, …

2. Scraping the Wikipedia Table

2a. Foreign-Exchange Reserves

I next applied the same rvest workflow to Wikipedia’s table of foreign-exchange reserves. Unlike the Books to Scrape sandbox, this was a real-world table with a more complicated structure.

wiki_url <- paste0("https://en.wikipedia.org/wiki/",
                   "List_of_countries_by_foreign-exchange_reserves")

wiki_pg <- read_html(wiki_url)

## Never assume the table you want is [[1]]. Look first.
tabs <- wiki_pg |> html_elements("table.wikitable")
length(tabs)                      # how many candidates? --> 2
[1] 2
tabs |> html_table() |> lapply(\(x) dim(x))
[[1]]
[1] 198   8

[[2]]
[1] 28 13
reserves_raw <- tabs[[1]] |> html_table()
glimpse(reserves_raw)
Rows: 198
Columns: 8
$ `Country(as recognized by the UN)` <chr> "Country(as recognized by the UN)",…
$ Continent                          <chr> "Continent", "Continent", "Asia", "…
$ `Foreign exchange reserves`        <chr> "Including gold", "US$ million", "3…
$ `Foreign exchange reserves`        <chr> "Including gold", "Change", "4,299"…
$ `Foreign exchange reserves`        <chr> "Excluding gold", "US$ million", "3…
$ `Foreign exchange reserves`        <chr> "Excluding gold", "Change", "115,49…
$ `Last reporteddate`                <chr> "Last reporteddate", "Last reported…
$ Ref.                               <chr> "Ref.", "Ref.", "[3]", "[4]", "[5]"…
reserves <- reserves_raw |>
  janitor::clean_names() |>
  mutate(
    across(
      where(is.character),
      ~ .x |>
        str_remove_all("\\[.*?\\]") |>
        str_squish()
    )
  ) |>
  filter(
    !str_detect(
      country_as_recognized_by_the_un,
      regex("total|world", ignore_case = TRUE)
    )
  )
reserves <- reserves[-c(1, 2), ]
glimpse(reserves)
Rows: 196
Columns: 8
$ country_as_recognized_by_the_un <chr> "China", "Japan", "Switzerland", "Indi…
$ continent                       <chr> "Asia", "Asia", "Europe", "Asia", "Eur…
$ foreign_exchange_reserves       <chr> "3,854,885", "1,287,470", "939,140", "…
$ foreign_exchange_reserves_2     <chr> "4,299", "18,400", "13,935", "45,290",…
$ foreign_exchange_reserves_3     <chr> "3,504,805", "1,259,248", "932,282", "…
$ foreign_exchange_reserves_4     <chr> "115,499", "16,230", "24,490", "56,790…
$ last_reporteddate               <chr> "1 August 2026", "30 June 2026", "30 J…
$ ref                             <chr> "", "", "", "", "", "", "", "", "", ""…
reserves <- reserves |>
  select(
    country_as_recognized_by_the_un,
    continent,
    foreign_exchange_reserves,
    last_reporteddate
  )
reserves <- reserves |>
  mutate(
    foreign_exchange_reserves = parse_number(foreign_exchange_reserves)
  )
head(reserves)
# A tibble: 6 × 4
  country_as_recognized_by_…¹ continent foreign_exchange_res…² last_reporteddate
  <chr>                       <chr>                      <dbl> <chr>            
1 China                       Asia                     3854885 1 August 2026    
2 Japan                       Asia                     1287470 30 June 2026     
3 Switzerland                 Europe                    939140 30 June 2026     
4 India                       Asia                      785710 04 September 2026
5 Russia                      Europe/A…                 761200 21 August 2026   
6 Taiwan                      Asia                      597150 30 June 2026     
# ℹ abbreviated names: ¹​country_as_recognized_by_the_un,
#   ²​foreign_exchange_reserves

2b. A Second Wikipedia Table

To test whether the same code would work on a different table, I used the FIFA World Cup results table from Wikipedia.

fifa_url <- "https://en.wikipedia.org/wiki/FIFA_World_Cup"
fifa_pg <- read_html(fifa_url)
fifa_tabs <- fifa_pg |>
  html_elements("table.wikitable")

length(fifa_tabs)
[1] 7
fifa_tabs |>
  html_table() |>
  lapply(\(x) dim(x))
[[1]]
[1] 6 4

[[2]]
[1] 28 10

[[3]]
[1] 26  6

[[4]]
[1] 10  8

[[5]]
[1] 11  5

[[6]]
[1] 11  3

[[7]]
[1] 23  8
tabs |> html_table() |> lapply(\(x) dim(x))
[[1]]
[1] 198   8

[[2]]
[1] 28 13
lapply(fifa_tabs, head)
[[1]]
[[1]]$node
<pointer: 0xa38cf5c80>

[[1]]$doc
<pointer: 0xa3d5d5680>


[[2]]
[[2]]$node
<pointer: 0xa38cb0d80>

[[2]]$doc
<pointer: 0xa3d5d5680>


[[3]]
[[3]]$node
<pointer: 0xa3e47eb00>

[[3]]$doc
<pointer: 0xa3d5d5680>


[[4]]
[[4]]$node
<pointer: 0xa3e515600>

[[4]]$doc
<pointer: 0xa3d5d5680>


[[5]]
[[5]]$node
<pointer: 0xa3e552d00>

[[5]]$doc
<pointer: 0xa3d5d5680>


[[6]]
[[6]]$node
<pointer: 0xa3e577500>

[[6]]$doc
<pointer: 0xa3d5d5680>


[[7]]
[[7]]$node
<pointer: 0xa3e819e80>

[[7]]$doc
<pointer: 0xa3d5d5680>
fifa_results_raw <- fifa_tabs[[7]] |>
  html_table()
glimpse(fifa_results_raw)
Rows: 23
Columns: 8
$ `World Cup`               <chr> "1930 Uruguay", "1934 Italy", "1938 France",…
$ `Golden Ball`             <chr> "Not Awarded", "Not Awarded", "Not Awarded",…
$ `Golden Boot`             <chr> "Guillermo Stábile", "Oldřich Nejedlý", "Leô…
$ Goals                     <int> 8, 5, 7, 9, 11, 13, 4, 9, 10, 7, 6, 6, 6, 6,…
$ `Golden Glove`            <chr> "Not Awarded", "Not Awarded", "Not Awarded",…
$ `Clean sheets`            <chr> "N/A", "N/A", "N/A", "N/A", "N/A", "N/A", "N…
$ `FIFA Young Player Award` <chr> "Not Awarded", "Not Awarded", "Not Awarded",…
$ `FIFA Fair Play Trophy`   <chr> "Not Awarded", "Not Awarded", "Not Awarded",…
head(fifa_results_raw, 10)
# A tibble: 10 × 8
   `World Cup`   `Golden Ball` `Golden Boot` Goals `Golden Glove` `Clean sheets`
   <chr>         <chr>         <chr>         <int> <chr>          <chr>         
 1 1930 Uruguay  Not Awarded   Guillermo St…     8 Not Awarded    N/A           
 2 1934 Italy    Not Awarded   Oldřich Neje…     5 Not Awarded    N/A           
 3 1938 France   Not Awarded   Leônidas          7 Not Awarded    N/A           
 4 1950 Brazil   Not Awarded   Ademir            9 Not Awarded    N/A           
 5 1954 Switzer… Not Awarded   Sándor Kocsis    11 Not Awarded    N/A           
 6 1958 Sweden   Not Awarded   Just Fontaine    13 Not Awarded    N/A           
 7 1962 Chile    Not Awarded   Flórián Albe…     4 Not Awarded    N/A           
 8 1966 England  Not Awarded   Eusébio           9 Not Awarded    N/A           
 9 1970 Mexico   Not Awarded   Gerd Müller      10 Not Awarded    N/A           
10 1974 West Ge… Not Awarded   Grzegorz Lato     7 Not Awarded    N/A           
# ℹ 2 more variables: `FIFA Young Player Award` <chr>,
#   `FIFA Fair Play Trophy` <chr>
#fifa_results_raw |>
  #janitor::clean_names() |>
 # filter(
  #  !str_detect(
  #    country_as_recognized_by_the_un,
  #    regex("total|world", ignore_case = TRUE) ) ) 

The scraping step itself still worked, but the cleaning code from the foreign-exchange table did not transfer directly. The FX table contained columns such as country_as_recognized_by_the_un and a two-row header, while the FIFA table had columns such as world_cup, golden_ball, and golden_boot. Therefore, cleaning commands that referenced the FX-specific columns could not be reused unchanged.

WarningWhat broke?

The extraction worked, but the cleaning code was tied to the structure of the original table. The FIFA table had different column names and did not contain the same multi-row header or Total/World rows. This showed that scraping code is often dependent on the specific structure of a webpage.

2c. Cleaning Data

fifa_results <- fifa_results_raw |>
  janitor::clean_names()

fifa_results <- fifa_results |>
  mutate(
    year = readr::parse_number(world_cup)
  )
fifa_results <- fifa_results |>
  mutate(
    across(
      where(is.character),
      ~ .x |>
        str_remove_all("\\[.*?\\]") |>
        str_squish()
    )
  )

fifa_results <- fifa_results |>
  mutate(
    across(
      where(is.character),
      ~ .x |>
        str_remove_all("\\[.*?\\]") |>
        str_squish()
    )
  )

head(fifa_results)
# A tibble: 6 × 9
  world_cup        golden_ball golden_boot       goals golden_glove clean_sheets
  <chr>            <chr>       <chr>             <int> <chr>        <chr>       
1 1930 Uruguay     Not Awarded Guillermo Stábile     8 Not Awarded  N/A         
2 1934 Italy       Not Awarded Oldřich Nejedlý       5 Not Awarded  N/A         
3 1938 France      Not Awarded Leônidas              7 Not Awarded  N/A         
4 1950 Brazil      Not Awarded Ademir                9 Not Awarded  N/A         
5 1954 Switzerland Not Awarded Sándor Kocsis        11 Not Awarded  N/A         
6 1958 Sweden      Not Awarded Just Fontaine        13 Not Awarded  N/A         
# ℹ 3 more variables: fifa_young_player_award <chr>,
#   fifa_fair_play_trophy <chr>, year <dbl>

3. Now Do It the Right Way

Wikipedia was useful for practicing scraping, but the same information is available through a structured World Bank indicator. This provides an opportunity to compare a scraped source with an API-based source.

3a. World Bank API

library(WDI)

reserves_api <- WDI(
  indicator = "FI.RES.TOTL.CD",
  start = 2015,
  end = 2026,
  extra = TRUE
)

head(reserves_api)
      country iso2c iso3c year FI.RES.TOTL.CD status lastupdated
1 Afghanistan    AF   AFG 2025             NA         2026-07-13
2 Afghanistan    AF   AFG 2024             NA         2026-07-13
3 Afghanistan    AF   AFG 2023             NA         2026-07-13
4 Afghanistan    AF   AFG 2022             NA         2026-07-13
5 Afghanistan    AF   AFG 2021             NA         2026-07-13
6 Afghanistan    AF   AFG 2020     9748946327         2026-07-13
                                             region capital longitude latitude
1 Middle East, North Africa, Afghanistan & Pakistan   Kabul   69.1761  34.5228
2 Middle East, North Africa, Afghanistan & Pakistan   Kabul   69.1761  34.5228
3 Middle East, North Africa, Afghanistan & Pakistan   Kabul   69.1761  34.5228
4 Middle East, North Africa, Afghanistan & Pakistan   Kabul   69.1761  34.5228
5 Middle East, North Africa, Afghanistan & Pakistan   Kabul   69.1761  34.5228
6 Middle East, North Africa, Afghanistan & Pakistan   Kabul   69.1761  34.5228
      income lending
1 Low income     IDA
2 Low income     IDA
3 Low income     IDA
4 Low income     IDA
5 Low income     IDA
6 Low income     IDA

FI.RES.TOTL.CD measures total reserves, including gold, in current U.S. dollars.

3b. Comparison

library(lubridate)

Attaching package: 'lubridate'
The following objects are masked from 'package:base':

    date, intersect, setdiff, union
# Convert Wikipedia's reporting date from character to Date
reserves_compare <- reserves |>
  mutate(
    last_reporteddate = dmy(last_reporteddate),
    
    # Wikipedia reports reserves in US$ million,
    # so convert to current US dollars
    wiki_reserves_usd = foreign_exchange_reserves * 1000000,
    
    # Extract the year from the reporting date
    wiki_year = year(last_reporteddate)
  ) |>
  select(
    country = country_as_recognized_by_the_un,
    last_reporteddate,
    wiki_year,
    wiki_reserves_usd
  )
Warning: There was 1 warning in `mutate()`.
ℹ In argument: `last_reporteddate = dmy(last_reporteddate)`.
Caused by warning:
!  1 failed to parse.
head(reserves_compare)
# A tibble: 6 × 4
  country     last_reporteddate wiki_year wiki_reserves_usd
  <chr>       <date>                <dbl>             <dbl>
1 China       2026-08-01             2026     3854885000000
2 Japan       2026-06-30             2026     1287470000000
3 Switzerland 2026-06-30             2026      939140000000
4 India       2026-09-04             2026      785710000000
5 Russia      2026-08-21             2026      761200000000
6 Taiwan      2026-06-30             2026      597150000000
## Match each country to the World Bank value for the same year ---------------

reserves_comparison <- reserves_compare |>
  left_join(
    reserves_api |>
      select(
        country,
        year,
        world_bank_reserves_usd = FI.RES.TOTL.CD,
        world_bank_last_updated = lastupdated
      ),
    by = c("country" = "country", "wiki_year" = "year")
  ) |>
  mutate(
    difference_usd = wiki_reserves_usd - world_bank_reserves_usd,
    difference_percent =
      (difference_usd / world_bank_reserves_usd) * 100
  ) |>
  arrange(desc(abs(difference_usd)))
##View the comparison table ---------------------------------------------------

reserves_comparison <- reserves_compare |>
  left_join(
    reserves_api |>
      select(
        country,
        wb_year = year,
        world_bank_reserves_usd = FI.RES.TOTL.CD,
        world_bank_last_updated = lastupdated
      ),
    by = c("country" = "country", "wiki_year" = "wb_year")
  ) |>
  mutate(
    difference_usd = wiki_reserves_usd - world_bank_reserves_usd,
    difference_percent =
      (difference_usd / world_bank_reserves_usd) * 100
  ) |>
  select(
    country,
    wiki_date = last_reporteddate,
    wiki_year,
    wiki_reserves_usd,
    world_bank_reserves_usd,
    world_bank_last_updated,
    difference_usd,
    difference_percent
  ) |>
  arrange(desc(abs(difference_percent)))
reserves_comparison |>
  select(
    country,
    wiki_date,
    wiki_year,
    wiki_reserves_usd,
    world_bank_reserves_usd,
    difference_usd,
    difference_percent
  ) |>
  head(20) |>
  mutate(
    wiki_reserves_usd = scales::comma(wiki_reserves_usd),
    world_bank_reserves_usd = scales::comma(world_bank_reserves_usd),
    difference_usd = scales::comma(difference_usd),
    difference_percent = paste0(round(difference_percent, 1), "%")
  ) |>
  knitr::kable(
    format = "html",
    col.names = c(
      "Country",
      "Wikipedia Date",
      "Year",
      "Wikipedia Reserves (US$)",
      "World Bank Reserves (US$)",
      "Difference (US$)",
      "Difference (%)"
    ),
    caption = "Comparison of Foreign-Exchange Reserves"
  )
Comparison of Foreign-Exchange Reserves
Country Wikipedia Date Year Wikipedia Reserves (US$) World Bank Reserves (US$) Difference (US$) Difference (%)
South Sudan 2024-03-02 2024 80,000,000 15,961,778 64,038,222 401.2%
Gabon 2024-03-01 2024 1,377,000,000 638,788,434 738,211,566 115.6%
North Macedonia 2025-08-31 2025 4,000,000 5,803,955,794 -5,799,955,794 -99.9%
Equatorial Guinea 2024-03-01 2024 46,000,000 1,076,807,833 -1,030,807,833 -95.7%
Belgium 2024-03-01 2024 82,000,000,000 42,330,768,016 39,669,231,984 93.7%
Chad 2024-03-01 2024 143,000,000 1,547,604,929 -1,404,604,929 -90.8%
United States 2025-08-22 2025 253,767,000,000 1,385,447,111,477 -1,131,680,111,477 -81.7%
Kenya 2024-03-05 2024 2,490,000,000 10,066,607,177 -7,576,607,177 -75.3%
Greece 2024-03-01 2024 3,926,000,000 15,221,750,195 -11,295,750,195 -74.2%
Zimbabwe 2024-03-01 2024 159,000,000 484,972,805 -325,972,805 -67.2%
Samoa 2024-03-01 2024 188,000,000 507,740,281 -319,740,281 -63%
Suriname 2024-03-08 2024 647,000,000 1,632,317,813 -985,317,813 -60.4%
Luxembourg 2024-03-05 2024 1,119,000,000 2,789,187,310 -1,670,187,310 -59.9%
Tajikistan 2024-03-01 2024 1,482,000,000 3,630,709,354 -2,148,709,354 -59.2%
Netherlands 2024-03-01 2024 125,451,000,000 79,129,609,286 46,321,390,714 58.5%
Lebanon 2024-03-01 2024 14,738,000,000 33,301,277,510 -18,563,277,510 -55.7%
Finland 2024-03-01 2024 7,995,000,000 17,992,892,153 -9,997,892,153 -55.6%
Barbados 2024-03-01 2024 770,000,000 1,653,675,115 -883,675,115 -53.4%
Ecuador 2024-03-01 2024 3,305,000,000 6,907,743,955 -3,602,743,955 -52.2%
Botswana 2024-03-01 2024 5,080,000,000 3,455,661,776 1,624,338,224 47%

Findings:

The two sources do not always report the same values for the same country and calendar year. The differences appear particularly large for some countries, showing that matching only by country and year does not necessarily mean that the two sources represent the same reporting vintage.

For example, Wikipedia reports U.S. reserves of approximately $253.8 billion with a reporting date of August 22, 2025, while the World Bank dataset reports approximately $1.39 trillion for 2025 and was updated in July 2026. This illustrates why the reporting date and dataset vintage matter when comparing values from different sources.

reserves_compare |>
  filter(country == "United States")
# A tibble: 1 × 4
  country       last_reporteddate wiki_year wiki_reserves_usd
  <chr>         <date>                <dbl>             <dbl>
1 United States 2025-08-22             2025      253767000000
reserves_api |>
  filter(country == "United States") |>
  select(country, year, FI.RES.TOTL.CD, lastupdated)
         country year FI.RES.TOTL.CD lastupdated
1  United States 2025   1.385447e+12  2026-07-13
2  United States 2024   9.100365e+11  2026-07-13
3  United States 2023   7.734262e+11  2026-07-13
4  United States 2022   7.066442e+11  2026-07-13
5  United States 2021   7.161523e+11  2026-07-13
6  United States 2020   6.283697e+11  2026-07-13
7  United States 2019   5.167006e+11  2026-07-13
8  United States 2018   4.499071e+11  2026-07-13
9  United States 2017   4.512853e+11  2026-07-13
10 United States 2016   4.059424e+11  2026-07-13
11 United States 2015   3.837285e+11  2026-07-13

3c. When is scraping the right tool?

Scraping is the right tool when the information you need is publicly available online but there is no suitable API, downloadable dataset, or other structured way to access it. It should be used when it is the only practical way to obtain the data, while respecting the website’s robots.txt, terms of service, and limits on requests.

NoteMain lesson

- Wikipedia: edited by volunteers, sourced from many places, no schema,

no versioning, no uncertainty, changes without notice

- World Bank: one compiler, documented methodology, stable indicator code

Scraping was the wrong tool here. Knowing when it is the only tool is the

skill this course is actually teaching.

4. Data Plan: Measuring Teacher Shortages in Dallas County ISDs

Research Question

Do administrative workforce indicators and online teacher vacancy records identify similar staffing challenges across public school districts in Dallas County?

This project will compare standardized administrative workforce indicators from the Texas Education Agency (TEA) with independently collected teacher vacancy records from district recruitment websites. The goal is to examine whether different approaches to measuring teacher shortages produce similar assessments of district staffing conditions.

4a. Data Sources

  • Primary unit of analysis: Independent school district (ISD)

  • Raw web observation: Individual job posting

  • Study area: Dallas County

Source What I would collect Why use it?
Texas Education Agency Teacher attrition, retention, employment, certification, enrollment, and other workforce indicators Standardized administrative data across Texas districts
District recruitment websites Job title, subject/position, posting date, district Direct measure of positions districts are actively advertising
NCES Contextual information such as teacher workforce and compensation Additional context if needed
Texas Schools Project Additional education data if needed Potential supplemental source

TEA provides standardized administrative measures of the teacher workforce, making it useful for comparing districts. District job postings measure a different aspect of staffing conditions: positions that districts are actively seeking to fill. Comparing these sources allows the project to examine whether administrative indicators and vacancy data tell the same story.

TipWhy use two measures?

Administrative data: What happened to the teacher workforce?

Vacancy data: What positions are districts currently trying to fill?

The comparison is the main research contribution of the data collection strategy.

4b. robots.txt and Terms of Service

We will check each recruitment site’s robots.txt file before collecting data and will review the Terms of Service where available. We will only automate collection from publicly accessible pages when the relevant path is permitted, and we will not bypass access restrictions, authentication, or other technical protections.

Our initial checks show that several district sites permit access to their public job-listing pages, while Highland Park ISD’s recruitment site currently disallows scraping. Sites with unclear or timed-out robots.txt checks will be verified before collection. If automated scraping is not permitted, we will exclude the site from automated collection or use manual collection if allowed.

District Recruitment platform Site type robots.txt / access
Irving ISD SchoolSpring Dynamic Path allowed
Dallas ISD Taleo Dynamic paths_allowed() returned TRUE
Plano ISD Applitrack Dynamic Relevant paths appear permitted
Richardson ISD Applitrack Dynamic Relevant paths appear permitted
Highland Park ISD TX Ed Job Network Dynamic Scraping disallowed
Carrollton-Farmers Branch ISD Applitrack Dynamic To be checked
Cedar Hill ISD District/Skyward-style system TBD Requires further inspection
Coppell ISD Winocular Dynamic Path allowed
DeSoto ISD Teams Dynamic Path allowed
Duncanville ISD Skyward Dynamic robots.txt retrieval worked; path check timed out
Garland ISD Applitrack Dynamic robots.txt retrieval worked; path check had an error
Grand Prairie ISD Skyward Dynamic robots.txt retrieval worked; path check timed out
Lancaster ISD Applitrack Dynamic robots.txt retrieval worked; path check had an error
Mesquite ISD Applitrack Dynamic To be checked
Sunnyvale ISD Applitrack Dynamic URL/platform needs verification

Highland Park ISD is currently not suitable for automated scraping because the recruitment site’s robots.txt disallows scraping. I would exclude it from automated collection rather than attempt to bypass the restriction.

4c. Whether an API exists

TEA

TEA data will be collected through available administrative datasets or structured dashboard data rather than scraped from the website.

District vacancies

The district recruitment systems do not share one common API because districts use different vendors and platforms. Where a suitable API or downloadable dataset is unavailable, permitted public vacancy pages may be collected through automated or manual methods, such as scraping, depending on the site’s structure. However, more research is being done for the most ethical way of obtaining the sites’ information.

4d. Volume, rate limiting, and cache

The project will collect a relatively small number of district vacancy pages rather than continuously crawling entire websites. Requests will be rate-limited and spaced out, and previously collected pages and extracted records will be stored locally so that pages do not need to be repeatedly requested. Each observation will include a retrieval date to distinguish the collection date from the original job posting date.

I will initially test the method on a small subset of districts before scaling it to the full study area.

4e. What breaks when the page changes

Potential change What could happen How I would detect it
HTML structure changes Scraper returns empty or incomplete data Check number of records and expected columns
Recruitment platform changes URL or selectors stop working Test each district before collection
Job fields change Posting date/title no longer extracted Validate expected fields
Site becomes more dynamic read_html() no longer captures listings Compare rendered page with extracted HTML
Robots.txt changes Automated collection may no longer be permitted Re-check robots.txt before collection
Old postings disappear Historical vacancy counts become incomplete Record retrieval date and archive collected results

Unexpected changes in the number of records or missing expected fields would trigger manual inspection rather than allowing the scraper to silently produce an incomplete dataset.