<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<title>pablo garcía-guzmán</title>
	<subtitle>essays on economics, data, and public policy</subtitle>
	<link href="https://pablogguz.github.io/feed.xml" rel="self" type="application/atom+xml"/>
    <link href="https://pablogguz.github.io/"/>
	<updated>2025-10-09T00:00:00+00:00</updated>
	<author><name>Pablo García-Guzmán</name></author>
	<id>https://pablogguz.github.io/feed.xml</id>
	<entry xml:lang="en">
		<title>No country for young men: age-earnings profiles by birth cohort and sex in Spain</title>
		<published>2025-10-09T00:00:00+00:00</published>
		<updated>2025-10-09T00:00:00+00:00</updated>
		<link href="https://pablogguz.github.io/blog/wage-cohort-esp/" type="text/html"/>
		<id>https://pablogguz.github.io/blog/wage-cohort-esp/</id>
		<content type="html">&lt;img src=&quot;&#x2F;img&#x2F;nocountry.png&quot; width=&quot;500&quot;&#x2F;&gt;
&lt;br&gt;
&lt;h1 id=&quot;introduction&quot;&gt;Introduction&lt;a class=&quot;zola-anchor&quot; href=&quot;#introduction&quot; aria-label=&quot;Anchor link for: introduction&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;Spain&#x27;s younger generations are not enjoying higher wages than their parents did at the same age. In fact, &lt;strong&gt;each successive male cohort since the 1960s appears to do progressively worse&lt;&#x2F;strong&gt;: men born in the 1990s, now in their mid-30s, earn about 9% less in real terms than men born in the 1970s at comparable ages.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Women&#x27;s wage trajectories tell a different story&lt;&#x2F;strong&gt;: recent cohorts have achieved parity with earlier generations and appear on track to surpass them. While matching or slightly exceeding wage levels from decades ago is hardly cause for celebration, it stands in marked contrast to the persistent declines younger men have experienced.&lt;&#x2F;p&gt;
&lt;p&gt;This post documents an attempt to construct age-earnings profiles by birth cohort in Spain. My approach accounts for methodological breaks in the underlying survey data and allows me to track different generations across their working lives.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;data&quot;&gt;Data&lt;a class=&quot;zola-anchor&quot; href=&quot;#data&quot; aria-label=&quot;Anchor link for: data&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;The main analysis draws on the following datasets, all publicly available:&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.ine.es&#x2F;jaxiT3&#x2F;Tabla.htm?t=28189&amp;amp;L=0&quot;&gt;Structure of Earnings Survey&lt;&#x2F;a&gt;&lt;&#x2F;strong&gt;: Estimates of average gross wages by age group and sex. My analysis includes the annual surveys from 2004-07 and 2008-23, as well as the quadrennial survey from 2002. Ages are reported in five-year bands (20-24, 25-29, ..., 55-59). These surveys are conducted by Spain&#x27;s National Statistical Office (&lt;em&gt;INE&lt;&#x2F;em&gt;) and cover salaried employees in industry, construction, and services (typically NACE Rev. 2 sections B to S), excluding agriculture, domestic services, self-employed workers, and some public administration positions.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;ec.europa.eu&#x2F;eurostat&#x2F;databrowser&#x2F;&#x2F;product&#x2F;view&#x2F;NAMA_10_GDP&quot;&gt;Eurostat National Accounts (nama_10_gdp)&lt;&#x2F;a&gt;&lt;&#x2F;strong&gt;: I use the household final consumption expenditure deflator (P31_S14_S15) to convert all wages to constant 2023 EUR.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.ine.es&#x2F;dyngs&#x2F;INEbase&#x2F;operacion.htm?c=Estadistica_C&amp;amp;cid=1254736176918&amp;amp;menu=resultados&amp;amp;idp=1254735976595#_tabs-1254736195129&quot;&gt;Labour Force Survey (EPA)&lt;&#x2F;a&gt;&lt;&#x2F;strong&gt;: I use data on sectoral employment shares to account for some methodological changes in the Structure of Earnings Survey. More on this below.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.mites.gob.es&#x2F;es&#x2F;estadisticas&#x2F;mercado_trabajo&#x2F;EMP&#x2F;welcome.htm&quot;&gt;&lt;em&gt;Estadística de Empresas Inscritas en la Seguridad Social&lt;&#x2F;em&gt; (Spanish Ministry of Social Security)&lt;&#x2F;a&gt;&lt;&#x2F;strong&gt;: I use data on employment counts by sector and firm size to account for some methodological changes in the Structure of Earnings Survey. More on this below.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;methodology&quot;&gt;Methodology&lt;a class=&quot;zola-anchor&quot; href=&quot;#methodology&quot; aria-label=&quot;Anchor link for: methodology&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;Data on wages is made available for five-year age bands, so the first step is to expand these five-year averages into single-year estimates while preserving the original means within each band. One way to do this is through a &lt;strong&gt;Lexis expansion&lt;&#x2F;strong&gt;.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;a-lexis-expansion-with-monotonic-splines&quot;&gt;A Lexis expansion with monotonic splines&lt;a class=&quot;zola-anchor&quot; href=&quot;#a-lexis-expansion-with-monotonic-splines&quot; aria-label=&quot;Anchor link for: a-lexis-expansion-with-monotonic-splines&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;In demography, a &lt;strong&gt;Lexis diagram&lt;&#x2F;strong&gt; is a graphical tool that represents demographic data across two dimensions: age (on the vertical axis) and time (on the horizontal axis). Diagonal lines in this diagram represent birth cohorts (i.e., groups of people born in the same year) as they age over time.&lt;&#x2F;p&gt;
&lt;p&gt;A &lt;strong&gt;Lexis expansion&lt;&#x2F;strong&gt; refers to the process of disaggregating data from coarser categories (like five-year age bands) into finer single-year units while maintaining consistency with the original aggregated data. The name comes from the fact that you are &quot;expanding&quot; the cells of a Lexis diagram from large blocks into smaller, more granular cells.&lt;&#x2F;p&gt;
&lt;p&gt;Specifically, for each calendar year $t$ and sex $s$, I proceed as follows:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Interpolate in log wages.&lt;&#x2F;strong&gt; I take the eight five-year bands $b\in \{ \text{20–24},\dots,\text{55–59} \}$ and place them at their midpoints $m_b\in\{22,27,\dots,57\}$. Then, I fit a &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Monotone_cubic_interpolation&quot;&gt;monotone cubic Hermite spline&lt;&#x2F;a&gt; to $\log \bar w_{t,b,s}$ over $m_b$, and predict $\tilde w_{t,a,s}$ for every integer age $a\in[20,59]$. This is a shape-preserving interpolator: where the observed mid-points are monotone, it prevents overshooting and spurious oscillations; where they are not monotone, it behaves like a constrained cubic Hermite. The reason for using a monotone Hermite spline instead of a regular cubic spline is that the latter can oscillate and overshoot band means between knots, while the former is explicitly designed to avoid those artifacts while still yielding a smooth profile.&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Ratio-adjust within each band to preserve published means.&lt;&#x2F;strong&gt; For each band $b$ with age set $A_b$ (e.g., $A_{\text{20–24}}={20,21...,24}$), compute&lt;&#x2F;p&gt;
&lt;p&gt;$$
r_{t,b,s} = \frac{\bar w_{t,b,s}}{\frac{1}{|A_b|}\sum_{a\in A_b}\tilde w_{t,a,s}}.
$$&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;Rescale ages inside the band:
$$
\hat w_{t,a,s} = r_{t,b,s}\times \tilde w_{t,a,s}\quad\text{for } a\in A_b.
$$
By construction,
$$
\frac{1}{|A_b|}\sum_{a\in A_b}\hat w_{t,a,s} = \bar w_{t,b,s}
$$
for every band $b$. This step “rakes” the interpolated curve so that the &lt;strong&gt;band averages match the official INE means exactly&lt;&#x2F;strong&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;Simple as that. Or is it?&lt;&#x2F;p&gt;
&lt;h2 id=&quot;not-so-fast&quot;&gt;Not so fast...&lt;a class=&quot;zola-anchor&quot; href=&quot;#not-so-fast&quot; aria-label=&quot;Anchor link for: not-so-fast&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;When working with survey data spanning many years, one must be cautious about potential changes in the data collection process and survey design. Such changes can lead to structural breaks in the series, making the data not comparable over time.&lt;&#x2F;p&gt;
&lt;p&gt;The Structure of Earnings Surveys in Spain have undergone several methodological changes over the years:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Surveys before 2008 &lt;strong&gt;excluded the public administration sector&lt;&#x2F;strong&gt; (NACE Rev. 2 section O &#x2F; NACE Rev. 1 section L). This sector tends to have a higher proportion of older workers and has higher average wages than the rest of the economy. Additionally, &lt;strong&gt;post-2008 only partially cover administration&lt;&#x2F;strong&gt;, including civil servants under the Social Security General Regime (&lt;em&gt;Régimen General&lt;&#x2F;em&gt;) but excluding those under special schemes such as MUFACE, ISFAS, and MUGEJU (&lt;em&gt;Clases Pasivas&lt;&#x2F;em&gt;). Since 2011, new civil servants are enrolled in the General Regime, so coverage has gradually increased over time. However, I could not find any readily available data on employment shares by sector and regime. For simplicity, and given that sensitivity tests suggest this has no impact on the results, the main analysis does not adjust for this partial coverage.&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#1&quot;&gt;1&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;&lt;&#x2F;li&gt;
&lt;li&gt;The quadrennial 2002 survey &lt;strong&gt;did not include firms with fewer than 10 employees&lt;&#x2F;strong&gt;, which tend to pay lower wages.&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;I address these issues as follows. I start by rescaling pre-2008 wages using an &lt;strong&gt;age-specific scaling factor&lt;&#x2F;strong&gt; that accounts for both the wage premium in public administration and its age-varying employment share:&lt;&#x2F;p&gt;
&lt;p&gt;$$
M_t^{(s)}(a) = 1 + s_{O,t}^{(s)}(a)\big(R_t^{(s)}-1\big)
$$&lt;&#x2F;p&gt;
&lt;p&gt;where (by year $(t)$, sex $(s)$, and age $(a)$):&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;$R_t^{(s)} = W_{O,t}^{(s)} &#x2F; W_{\text{rest},t}^{(s)}$ = &lt;strong&gt;overall wage ratio&lt;&#x2F;strong&gt; for public administration relative to the rest of sectors (averaged across all ages)&lt;&#x2F;li&gt;
&lt;li&gt;$s_{O,t}^{(s)}(a) = \dfrac{N_{O,t}^{(s)}(a)}{N_{O,t}^{(s)}(a) + N_{\text{rest},t}^{(s)}(a)}$ = &lt;strong&gt;age-specific employment share of O&lt;&#x2F;strong&gt; in B–S&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;The age variation in the scaling factor arises entirely from differences in sectoral employment composition across ages, rather than from age-varying wage premia. This approach leverages detailed employment breakdowns by age, sex and sector available in the &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.ine.es&#x2F;jaxiT3&#x2F;Tabla.htm?t=65986&amp;amp;L=0&quot;&gt;Labour Force Survey&lt;&#x2F;a&gt; (&lt;em&gt;EPA&lt;&#x2F;em&gt;) for 2008-2023, while using the overall wage ratio (which lacks the needed age-sex-sector disaggregation in the published data).&lt;&#x2F;p&gt;
&lt;p&gt;Since the adjustment must be applied &quot;out-of-sample&quot; for the pre-2008 period, the assumption is that the scaling factor is stable over time. &lt;a href=&quot;#fig_shares&quot;&gt;Figure 1&lt;&#x2F;a&gt; shows that this is indeed the case.&lt;&#x2F;p&gt;
&lt;figure id=&quot;fig_shares&quot;&gt;
  &lt;figcaption&gt;Figure 1: Scaling factors for public administration exclusion &lt;&#x2F;figcaption&gt;
   &lt;img src=&quot;&#x2F;img&#x2F;m_sector_by_age.png&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; border: 0px none;&quot;&gt;
   &lt;p class=&quot;figure-note&quot;&gt;Source: Structure of Earnings Survey and Labour Force Survey (Spanish Statistical Office) and author&#x27;s calculations. Note: The scaling factor $M^{(s)}(a) = 1 + s_O^{(s)}(a)(R^{(s)} - 1)$ combines the age-specific employment share of public administration with the overall wage premium for that sector. &lt;&#x2F;p&gt;
&lt;&#x2F;figure&gt;
&lt;br&gt;
&lt;p&gt;To avoid picking an arbitrary year, I average the obtained scaling factors over the 2008-2023 period for each age group.&lt;&#x2F;p&gt;
&lt;p&gt;Similarly, I &lt;strong&gt;rescale pre-2004 wages to account for the exclusion of firms with fewer than 10 employees&lt;&#x2F;strong&gt;. These micro-firms tend to pay lower wages than larger establishments, so pre-2004 surveys systematically overstate average wages. The scaling factor follows analogous logic:&lt;&#x2F;p&gt;
&lt;p&gt;$$
M_{t}^{\text{size},(s)} = 1 + s_{\text{micro},t}^{(s)}\big(r^{(s)}-1\big)
$$&lt;&#x2F;p&gt;
&lt;p&gt;where (by sex $(s)$):&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;$r^{(s)} = W_{\text{micro}}^{(s)} &#x2F; W_{10+}^{(s)}$ = wage ratio for micro-firms (1-9 employees) relative to firms with 10+ employees&lt;&#x2F;li&gt;
&lt;li&gt;$s_{\text{micro},t}^{(s)}$ = sex-specific employment share of micro-firms across covered sectors&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;Computing these components requires combining multiple data sources. For the employment share $s_{\text{micro}}^{(s)}$, I proceed in two steps. First, I obtain sector-specific employment shares by firm size from the Spanish Ministry of Social Security (&lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.mites.gob.es&#x2F;es&#x2F;estadisticas&#x2F;mercado_trabajo&#x2F;EMP&#x2F;welcome.htm&quot;&gt;&lt;em&gt;Estadística de Empresas Inscritas en la Seguridad Social&lt;&#x2F;em&gt;&lt;&#x2F;a&gt;). This dataset provides annual counts of employment by firm size categories for all NACE Rev. 2 sectors covered in the 2008-23 Structure of Earnings Survey (except public administration), which I use to compute sector-year specific micro-firm shares $s_{\text{micro},k,t}$ for each sector $k$.&lt;&#x2F;p&gt;
&lt;p&gt;Second, &lt;strong&gt;I weight these micro-firm shares by sex-specific sectoral employment shares&lt;&#x2F;strong&gt; from the Labour Force Survey:&lt;&#x2F;p&gt;
&lt;p&gt;$$
s_{\text{micro},t}^{(s)} = \sum_{k} \omega_{k,t}^{(s)} \cdot s_{\text{micro}, k, t}
$$&lt;&#x2F;p&gt;
&lt;p&gt;where $\omega_{k,t}^{(s)}$ is the share of sex $s$&#x27;s employment in sector $k$ at year $t$, computed only over the sectors B-S covered by the Structure of Earnings Survey.&lt;&#x2F;p&gt;
&lt;p&gt;For the wage ratio $r^{(s)}$, &lt;strong&gt;I rely on published tabulations from the 2006 Structure of Earnings Survey&lt;&#x2F;strong&gt; (see &lt;a href=&quot;#grafico38&quot;&gt;&lt;em&gt;Gráfico 38&lt;&#x2F;em&gt;&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.ine.es&#x2F;daco&#x2F;daco42&#x2F;salarial&#x2F;prinre06.pdf&quot;&gt;here&lt;&#x2F;a&gt;), which provides average wages by firm size and sex. The ratio is derived by solving:&lt;&#x2F;p&gt;
&lt;p&gt;$$
W_{\text{all}}^{(s)} = s_{\text{micro}}^{(s)} \cdot W_{\text{micro}}^{(s)} + (1 - s_{\text{micro}}^{(s)}) \cdot W_{10+}^{(s)}
$$&lt;&#x2F;p&gt;
&lt;p&gt;for $r^{(s)} = W_{\text{micro}}^{(s)} &#x2F; W_{10+}^{(s)}$.&lt;&#x2F;p&gt;
&lt;p&gt;Since the firm size adjustment must be applied to the 2002 survey, the key assumption is again that this scaling factor remains stable over time. The resulting scaling factors are shown in &lt;a href=&quot;#fig_shares_size&quot;&gt;Figure 2&lt;&#x2F;a&gt;, which displays the evolution of $M_{t}^{\text{size},(s)}$ over time by sex, and confirms its stability. These adjustments are more substantial than the public administration correction, reflecting the relatively large share of employment in small firms and their lower average wages. For the analysis, I average the obtained scaling factors over the entire 2013-2023 period for each sex.&lt;&#x2F;p&gt;
&lt;figure id=&quot;fig_shares_size&quot;&gt;
  &lt;figcaption&gt;Figure 2: Scaling factors for micro-firms exclusion &lt;&#x2F;figcaption&gt;
   &lt;img src=&quot;&#x2F;img&#x2F;m_size_by_sex.png&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; border: 0px none;&quot;&gt;
   &lt;p class=&quot;figure-note&quot;&gt;Source: Spanish Ministry of Social Security (Estadística de Empresas Inscritas en la Seguridad Social), Spanish Statistical Office (Labour Force Survey and Structure of Earnings Survey) and author&#x27;s calculations. Note: The scaling factor $M^{\text{size},(s)} = 1 + s_{\text{micro}}^{(s)}(r^{(s)} - 1)$ combines the sex-specific employment share of micro-firms (1-9 employees) with the wage penalty these firms pay relative to larger establishments.&lt;&#x2F;p&gt;
&lt;&#x2F;figure&gt;
&lt;!-- &lt;div id=&quot;table_firmsize&quot;&gt;
&lt;figcaption&gt;Table 2: Scaling factors for micro-firms exclusion&lt;&#x2F;figcaption&gt;

| Sex | Micro-firm share ($s_{\text{micro}}$) | Wage ratio ($r$) | Scaling factor ($M^{\text{size}}$) |
|-----|---------------------------------------|------------------|-------------------------------------|
| Both sexes | 21.8% | 0.698 | 0.934 |
| Males | 22.6% | 0.696 | 0.931 |
| Females | 20.8% | 0.698 | 0.937 |

&lt;p class=&quot;figure-note&quot;&gt;Source: Spanish Ministry of Social Security (Estadística de Empresas Inscritas en la Seguridad Social), Spanish Statistical Office (Labour Force Survey and Structure of Earnings Survey) and author&#x27;s calculations.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt; --&gt;
&lt;p&gt;With these two adjustments in place, we are ready to proceed with the Lexis expansion as described above.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;cohorts-and-indexing&quot;&gt;Cohorts and indexing&lt;a class=&quot;zola-anchor&quot; href=&quot;#cohorts-and-indexing&quot; aria-label=&quot;Anchor link for: cohorts-and-indexing&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h3&gt;
&lt;p&gt;With single-year ages, cohorts are straightforward:&lt;&#x2F;p&gt;
&lt;p&gt;$$c = t - a$$&lt;&#x2F;p&gt;
&lt;p&gt;I group individual cohorts into decades (1950s, 1960s, etc.). For each cohort decade $d$, the average wage at age $a$ is just:&lt;&#x2F;p&gt;
&lt;p&gt;$$\bar{w}(a, d) = \frac{1}{|C_d|} \sum_{c \in C_d} w^{Lexis}(c, a)$$&lt;&#x2F;p&gt;
&lt;p&gt;I index everything relative to the 1960s cohort at age 40 separately by sex. That is, an index of 80 at age 30 means that the cohort earns 80% of what the 1960s cohort of the same sex earned at age 40.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;results-and-discussion&quot;&gt;Results and discussion&lt;a class=&quot;zola-anchor&quot; href=&quot;#results-and-discussion&quot; aria-label=&quot;Anchor link for: results-and-discussion&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;The results are shown in &lt;a href=&quot;#fig3&quot;&gt;Figure 3&lt;&#x2F;a&gt;. &lt;strong&gt;For males, the profile suggests that each successive cohort has experienced weaker wage growth throughout their working lives&lt;&#x2F;strong&gt;. By age 35, men born in the 1980s earned almost 10% less in real terms than the 1960s cohort did at the same age. The 1990s cohort has failed to catch up and seems to be on a similar trajectory to those born in the 1980s, with earnings about 9% lower than what the 1970s cohort earned at comparable ages.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;For females, the picture is different but hardly encouraging&lt;&#x2F;strong&gt;. Recent cohorts show signs of convergence, with women born in the 1990s reaching parity with the 1970s cohort by their late-20s and outperforming the 1980s cohort by their early-30s. In this context, however, &quot;convergence&quot; means matching or slightly exceeding wage levels from decades ago. I will let the reader decide whether this warrants celebration.&lt;&#x2F;p&gt;
&lt;figure id=&quot;fig3&quot;&gt;
  &lt;figcaption&gt;Figure 3 &lt;&#x2F;figcaption&gt;
  &lt;img src=&quot;&#x2F;img&#x2F;cohort_wage_trajectories_by_sex_indexed.png&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; border: 0px none;&quot;&gt;&lt;&#x2F;img&gt;
&lt;&#x2F;figure&gt;
&lt;p&gt;As a result, &lt;strong&gt;the gender wage gap has narrowed substantially across cohorts&lt;&#x2F;strong&gt;. &lt;a href=&quot;#fig4&quot;&gt;Figure 4&lt;&#x2F;a&gt; shows the female-to-male wage ratio by cohort and age. Women born in the 1990s earn 88 cents for every euro their male counterparts make at ages 25 and above, up from 83 cents for the 1970s cohort at similar ages. In other words, the gender wage gap has shrunk by roughly 30% in two decades.&lt;&#x2F;p&gt;
&lt;figure id=&quot;fig4&quot;&gt;
  &lt;figcaption&gt;Figure 4 &lt;&#x2F;figcaption&gt;
  &lt;img src=&quot;&#x2F;img&#x2F;gender_wage_gap_by_cohort.png&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; border: 0px none;&quot;&gt;&lt;&#x2F;img&gt;
&lt;&#x2F;figure&gt;
&lt;p&gt;Discussing the root causes of this divergence is beyond the scope of this post, but it would be an interesting avenue for future work. I speculate on a couple of mechanisms below.&lt;&#x2F;p&gt;
&lt;p&gt;First, one could think about &lt;strong&gt;structural change&lt;&#x2F;strong&gt;. The 2008 financial crisis and subsequent Eurozone crisis disproportionately affected sectors of the economy that were dominated by male workers, particularly in construction and manufacturing (see &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.journals.uchicago.edu&#x2F;doi&#x2F;full&#x2F;10.1086&#x2F;718660&quot;&gt;here&lt;&#x2F;a&gt; and &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;pmc.ncbi.nlm.nih.gov&#x2F;articles&#x2F;PMC8488327&#x2F;?utm_source=chatgpt.com&quot;&gt;here&lt;&#x2F;a&gt;). These sectors never fully recovered, and if labour supply was absorbed by low-skill jobs in service sectors like retail and hospitality, this could have contributed to the observed wage stagnation for younger cohorts. Descriptively, one indeed sees that manufacturing employment has been in steady decline since the early 1990s, while construction employment collapsed after 2008 and has only partially recovered since then (see &lt;a href=&quot;#fig5&quot;&gt;Figure 5&lt;&#x2F;a&gt;). Employment in accommodation and food services has increased in relative terms, but there does not seem to be any sex-specific pattern. A &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.sciencedirect.com&#x2F;science&#x2F;article&#x2F;abs&#x2F;pii&#x2F;S0305750X13002246&quot;&gt;Rodrik-style within&#x2F;between decomposition&lt;&#x2F;a&gt; of wage growth by sex could tell us how much of the cohort gap comes from within-sector wage changes vs. shifts across sectors.&lt;&#x2F;p&gt;
&lt;figure id=&quot;fig5&quot;&gt;
  &lt;figcaption&gt;Figure 5 &lt;&#x2F;figcaption&gt;
  &lt;img src=&quot;&#x2F;img&#x2F;employment_shares_by_sector.png&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; border: 0px none;&quot;&gt;
&lt;&#x2F;figure&gt;
&lt;p&gt;Second, the analysis in this post focuses on average wages, but &lt;strong&gt;cohort gaps might reflect differences in hours worked and job intensity&lt;&#x2F;strong&gt;. In Spain, output per hour &lt;em&gt;has&lt;&#x2F;em&gt; increased over the last decades, but these improvements have not translated into higher average wages – instead, they have been absorbed by reductions in hours worked (see &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;cincodias.elpais.com&#x2F;economia&#x2F;2025-06-23&#x2F;el-enigma-de-la-productividad-espanola-y-si-la-estabamos-midiendo-mal.html#?prm=copy_link&quot;&gt;here&lt;&#x2F;a&gt;). The gender divergence could also reflect a disproportionate reduction in male hours among younger cohorts.&lt;&#x2F;p&gt;
&lt;p&gt;Whatever the mechanisms, these numbers tell a grim story. The absence or even reversal of improvements in labour market outcomes across generations represents the collapse of an implicit contract that held for decades: work hard and society will reward you with opportunities for upward mobility. That promise seems to be evaporating. A dysfunctional labour market, housing costs that previous generations never faced relative to their income, and the fiscal burden of an ageing population have created a perfect storm. Spain&#x27;s younger generations have inherited an economy that has turned against them.&lt;&#x2F;p&gt;
&lt;p&gt;Contemporary Spain is no country for young men. And increasingly, no country for the young.&lt;&#x2F;p&gt;
&lt;hr &#x2F;&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;1&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;1&lt;&#x2F;sup&gt;
&lt;p&gt;I tested the sensitivity of the results by including an adjustment for the partial coverage of public administration when computing the sectoral employment shares, assuming a range of plausible coverage rates of the General Regime. The results are virtually unchanged.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
</content>
	</entry>
	<entry xml:lang="en">
		<title>A demographic projection of Spain&#x27;s native-born population: quantifying the stock of second-generation immigrants</title>
		<published>2025-05-08T00:00:00+00:00</published>
		<updated>2025-05-08T00:00:00+00:00</updated>
		<link href="https://pablogguz.github.io/blog/pop-projections-esp/" type="text/html"/>
		<id>https://pablogguz.github.io/blog/pop-projections-esp/</id>
		<content type="html">&lt;img src=&quot;&#x2F;img&#x2F;bull.png&quot; width=&quot;500&quot;&#x2F;&gt;
&lt;br&gt;
&lt;h1 id=&quot;introduction&quot;&gt;Introduction&lt;a class=&quot;zola-anchor&quot; href=&quot;#introduction&quot; aria-label=&quot;Anchor link for: introduction&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;Spain&#x27;s demographic structure is changing — and fast. In 1998, the stock of foreign-born residents was just 2.9% of the total population. By 2024, this figure had skyrocketed to 18.2%, with nearly one in five residents born outside Spain. Among the prime working-age population (ages 25-54), the share of foreign-born individuals has reached 26.8% nationally, and in provinces hosting major urban and economic centers (like Madrid or Barcelona) it is as high as ~35%. That is more than one in three people! &lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#1&quot;&gt;1&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;&lt;&#x2F;p&gt;
&lt;p&gt;This demographic shift extends beyond the first generation of migrants. The children of immigrants — often referred to as &quot;second-generation immigrants&quot; or, more precisely, native-born individuals with foreign-born parents — represent an increasingly important demographic group whose size and characteristics will shape Spain&#x27;s social and economic landscape in the coming decades.&lt;&#x2F;p&gt;
&lt;p&gt;Unfortunately, there is no readily available data on the number of second-generation immigrants over time in Spain. To bridge this gap, &lt;strong&gt;in this post I present a methodological framework for projecting the share of native-born individuals with foreign-born mothers across Spanish provinces from 2024 to 2039&lt;&#x2F;strong&gt;. To the best of my knowledge, this is the first attempt to estimate and project the size of this population group into the future. The methodology integrates multiple data sources from Spain&#x27;s National Statistical Office (&lt;em&gt;Instituto Nacional de Estadística&lt;&#x2F;em&gt;, hereafter INE) and employs simple, fully-reproducible modeling techniques to produce province-level projections of this increasingly important population segment.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;data&quot;&gt;Data&lt;a class=&quot;zola-anchor&quot; href=&quot;#data&quot; aria-label=&quot;Anchor link for: data&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;The projection methodology I propose draws on multiple datasets. All data is publicly available and can be accessed through INE&#x27;s website:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.ine.es&#x2F;dyngs&#x2F;INEbase&#x2F;es&#x2F;operacion.htm?c=Estadistica_C&amp;amp;cid=1254736176953&amp;amp;menu=ultiDatos&amp;amp;idp=1254735572981&quot;&gt;&lt;strong&gt;Population projections (2024-2039)&lt;&#x2F;strong&gt;&lt;&#x2F;a&gt;: This is the backbone of the analysis. The projections (elaborated by INE) provide estimates of Spain&#x27;s future population broken down by basic demographic characteristics following cohort component methods, which are the international standard for population projections. The projections also include mortality indicators, which will be used to estimate survival probabilities for each age group. All the projections used in the analysis correspond to the &quot;central&quot; (baseline) scenario. One way to think about this baseline scenario is as a &quot;business as usual&quot; projection, which assumes that current demographic trends (birth rates, death rates, and migration patterns) will continue into the future along their present trajectory.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.ine.es&#x2F;dyngs&#x2F;INEbase&#x2F;es&#x2F;operacion.htm?c=Estadistica_C&amp;amp;cid=1254736177007&amp;amp;menu=ultiDatos&amp;amp;idp=1254735573002&quot;&gt;&lt;strong&gt;Birth statistics&lt;&#x2F;strong&gt;&lt;&#x2F;a&gt;: The births microdata allows me to establish the baseline proportion of births to foreign-born mothers in each province and provides historical data for modeling the relationship between foreign-born female population and births to foreign-born mothers.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.ine.es&#x2F;dyngs&#x2F;INEbase&#x2F;en&#x2F;operacion.htm?c=Estadistica_C&amp;amp;cid=1254736177092&amp;amp;menu=ultiDatos&amp;amp;idp=1254735572981&quot;&gt;&lt;strong&gt;Survey on Essential Characteristics of Population and Housing (ECEPOV-2021)&lt;&#x2F;strong&gt;&lt;&#x2F;a&gt;: This household survey, conducted by INE in 2021, complements the 2021 census by providing information not available in administrative registers — in particular, the ECEPOV-2021 will allow me to calculate the existing share of native-born individuals with foreign-born mothers by age and province in 2021. This will provide the baseline structure for the projections.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h1 id=&quot;methodology&quot;&gt;Methodology&lt;a class=&quot;zola-anchor&quot; href=&quot;#methodology&quot; aria-label=&quot;Anchor link for: methodology&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;h2 id=&quot;the-fundamentals&quot;&gt;The fundamentals&lt;a class=&quot;zola-anchor&quot; href=&quot;#the-fundamentals&quot; aria-label=&quot;Anchor link for: the-fundamentals&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;Let&#x27;s start with the basics. The classic equation that demographers use to track population changes (in its discrete form) is quite simple:&lt;&#x2F;p&gt;
&lt;p&gt;$$N_{t+1} = N_t + B_t - D_t + I_t - E_t$$&lt;&#x2F;p&gt;
&lt;p&gt;where $N_t$ is the population at time $t$, $B_t$ is births, $D_t$ is deaths, $I_t$ is immigration, and $E_t$ is emigration. This equation captures the fundamental flows that shape population: we start with last year&#x27;s population, add new births, subtract deaths, and account for people moving in and out.&lt;&#x2F;p&gt;
&lt;p&gt;Let me now write down a slightly more refined version of the equation above that accounts for age structure:&lt;&#x2F;p&gt;
&lt;p&gt;$$N(a+1,t+1) = N(a,t) \times s(a,t) + M(a,t)$$&lt;&#x2F;p&gt;
&lt;p&gt;Here, $N(a,t)$ represents the population of age $a$ at time $t$, while $s(a,t)$ is the survival probability (the chance someone of age $a$ survives to age $a+1$), and $M(a,t)$ captures net migration (immigration minus emigration) of that particular age group. For example, to find next year&#x27;s 30-year-olds, take this year&#x27;s 29-year-olds, calculate how many will survive to age 30, and add any net migration of 30-year-olds.&lt;&#x2F;p&gt;
&lt;p&gt;Now, we&#x27;re specifically interested in the stock of second-generation immigrants. &lt;strong&gt;For simplicity, I will define this group as native-born individuals with foreign-born mothers (hereafter NBF)&lt;&#x2F;strong&gt;. We can then decompose the total population into three mutually exclusive groups:&lt;&#x2F;p&gt;
&lt;p&gt;$$N(a,t) = NBF(a,t) + NBNM(a,t) + FB(a,t)$$&lt;&#x2F;p&gt;
&lt;p&gt;where:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;$NBF(a,t)$ is the native-born population with foreign-born mothers&lt;&#x2F;li&gt;
&lt;li&gt;$NBNM(a,t)$ is the native-born population with native-born mothers&lt;&#x2F;li&gt;
&lt;li&gt;$FB(a,t)$ is the foreign-born population&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;Each component follows the same demographic principles but differs in how new members enter the population:&lt;&#x2F;p&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Population group&lt;&#x2F;th&gt;&lt;th&gt;Change equation&lt;&#x2F;th&gt;&lt;th&gt;Newborns&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;Native-born with foreign-born mothers (NBF)&lt;&#x2F;td&gt;&lt;td&gt;$NBF(a+1,t+1) = NBF(a,t) \times s(a,t) + M_{NBF}(a,t)$&lt;&#x2F;td&gt;&lt;td&gt;$NBF(0,t+1) = B(t) \times s_{infant}(p,t) \times \Phi_{FB}(t) + M_{NBF}(0,t)$&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Native-born with native-born mothers (NBNM)&lt;&#x2F;td&gt;&lt;td&gt;$NBNM(a+1,t+1) = NBNM(a,t) \times s(a,t) + M_{NBNM}(a,t)$&lt;&#x2F;td&gt;&lt;td&gt;$NBNM(0,t+1) = B(t) \times s_{infant}(p,t) \times (1 - \Phi_{FB}(t)) + M_{NBNM}(0,t)$&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Foreign-born population (FB)&lt;&#x2F;td&gt;&lt;td&gt;$FB(a+1,t+1) = FB(a,t) \times s(a,t) + M_{FB}(a,t)$&lt;&#x2F;td&gt;&lt;td&gt;$FB(0,t+1) = M_{FB}(0,t)$&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;p&gt;In the table above, $\Phi_{FB}(t)$ is the share of births to foreign-born mothers in year $t$, and $M_{NBF}(a,t)$, $M_{NBNM}(a,t)$, and $M_{FB}(a,t)$ are the migration flows for each population group.&lt;&#x2F;p&gt;
&lt;p&gt;Notice that $s_{infant}(p,t) \neq s(p,0,t)$. The intuition is the following. Demographic projections typically reference population stocks as of January 1st of each year, which means that the age 0 population at time $t+1$ consists of children born during year $t$ who survived to January 1st. Since births occur throughout year $t$, newborns are exposed to mortality risk for only part of the year on average. To account for this partial-year exposure, I use an adjusted infant survival probability that differs from standard age-specific survival rates:&lt;&#x2F;p&gt;
&lt;p&gt;$$s_{infant}(p,t) = 1 - 0.5 \times m(0,p,t)$$&lt;&#x2F;p&gt;
&lt;p&gt;where $m(0,p,t)$ is the mortality rate for age 0 in province $p$ during year $t$. The 0.5 multiplier reflects that, on average, newborns are exposed to mortality risk for approximately half a year before reaching the January 1st reference date.&lt;&#x2F;p&gt;
&lt;!-- In practical terms, I implement this as follows:
-  I project the NBF population forward using only births and deaths, $NBF_{projected}(a,t)$
-  I compare the total population projection with INE&#x27;s official projections (which includes migration): $M(t) = N_{official}(t) - N_{projected}(t)$

The NBF projection is therefore given by:

$$NBF_{adjusted}(a,t) = NBF_{projected}(a,t) \times \frac{N_{official}(a,t)}{N_{projected}(a,t)}$$ --&gt;
&lt;h2 id=&quot;modeling-births-to-foreign-born-mothers&quot;&gt;Modeling births to foreign-born mothers&lt;a class=&quot;zola-anchor&quot; href=&quot;#modeling-births-to-foreign-born-mothers&quot; aria-label=&quot;Anchor link for: modeling-births-to-foreign-born-mothers&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;To project future demographic patterns accurately, we need some assumption(s) on how the proportion of births to foreign-born mothers will evolve over time. In particular, this key parameter isn&#x27;t static but fluctuates in response to changes in the foreign-born female population. One simple way to think about it is using a &lt;strong&gt;linear regression model that links the share of foreign-born women in a province to the percentage of births to foreign-born mothers in that same province&lt;&#x2F;strong&gt;:&lt;&#x2F;p&gt;
&lt;p&gt;$$\Phi_{FB}(p,t) = \alpha_p + \beta \times \frac{\text{Foreign-born females} (p,t)}{\text{Total females} (p,t)} + \varepsilon_{p,t}$$&lt;&#x2F;p&gt;
&lt;p&gt;where:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;$\Phi_{FB}(p,t)$ is the share of births to foreign-born mothers in province $p$ at time $t$&lt;&#x2F;li&gt;
&lt;li&gt;$\alpha_p$ represents province fixed effects&lt;&#x2F;li&gt;
&lt;li&gt;$\beta$ is a global elasticity coefficient&lt;&#x2F;li&gt;
&lt;li&gt;$\frac{\text{Foreign-born females} (p,t)}{\text{Total females} (p,t)}$ is the proportion of women in province $p$ at time $t$ who were born abroad.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;I estimate the equation above using data from 2009 to 2023, and then apply the estimated coefficients to projected foreign-born population shares through 2039. In the out-of-sample exercise, the number of foreign-born females is calibrated to match the ratio of overall foreign-born population to the total population in each province and year. That is:&lt;&#x2F;p&gt;
&lt;p&gt;$$\frac{\text{Foreign-born females} (p,t)}{\text{Total females} (p,t)} = \frac{\text{Foreign-born population} (p,t)}{\text{Total population} (p,t)} = \gamma(p,t)$$&lt;&#x2F;p&gt;
&lt;p&gt;Ideally, one would use the foreign-born share among the female population of childbearing age (15-49) instead of the overall share of foreign-born women, but unfortunately INE&#x27;s projections do not provide this breakdown (which makes the out-of-sample prediction exercise unfeasible unless otherwise estimated). For simplicity, I use the overall share as a proxy.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;a-basic-births-and-deaths-projection-model&quot;&gt;A basic births-and-deaths projection model&lt;a class=&quot;zola-anchor&quot; href=&quot;#a-basic-births-and-deaths-projection-model&quot; aria-label=&quot;Anchor link for: a-basic-births-and-deaths-projection-model&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;Let&#x27;s abstract from migration for a second. Under a pure births-and-deaths framework, or each province $p$ and year $t$, the population of age $a$ in year $t+1$ is given by:&lt;&#x2F;p&gt;
&lt;p&gt;$$\hat{N}(p,a,t+1) = \underbrace{\hat{N}(p,a-1,t) \times s(p,a-1,t)}_{\text{Aging and survival of last year&#x27;s population}}$$&lt;&#x2F;p&gt;
&lt;p&gt;Similarly, for the native-born population with foreign-born mothers and native-born population with native-born mothers:&lt;&#x2F;p&gt;
&lt;p&gt;$$\widehat{NBF}(p,a,t+1) = \underbrace{\widehat{NBF}(p,a-1,t) \times s(p,a-1,t)}_{\text{Aging and survival of last year&#x27;s NBF population}}$$&lt;&#x2F;p&gt;
&lt;p&gt;$$\widehat{NBNM}(p,a,t+1) = \underbrace{\widehat{NBNM}(p,a-1,t) \times s(p,a-1,t)}_{\text{Aging and survival of last year&#x27;s NBNM population}}$$&lt;&#x2F;p&gt;
&lt;p&gt;And for newborns:&lt;&#x2F;p&gt;
&lt;p&gt;$$\widehat{NBF}(p,0,t+1) = B(p,t) \times s_{infant}(p,t) \times \widehat{\Phi_{FB}}(p,t)$$&lt;&#x2F;p&gt;
&lt;p&gt;$$\widehat{NBNM}(p,0,t+1) = B(p,t) \times s_{infant}(p,t)  \times (1-\widehat{\Phi_{FB}}(p,t))$$&lt;&#x2F;p&gt;
&lt;p&gt;where $\widehat{\Phi_{FB}}(p,t)$ is estimated using the predicted values from the regression model:&lt;&#x2F;p&gt;
&lt;p&gt;$$\widehat{\Phi_{FB}}(p,t)= \widehat{\alpha_p} + \widehat{\beta} \times \gamma(p,t)$$&lt;&#x2F;p&gt;
&lt;h2 id=&quot;estimating-the-initial-stock-of-nbf-and-nbnm-individuals&quot;&gt;Estimating the initial stock of NBF and NBNM individuals&lt;a class=&quot;zola-anchor&quot; href=&quot;#estimating-the-initial-stock-of-nbf-and-nbnm-individuals&quot; aria-label=&quot;Anchor link for: estimating-the-initial-stock-of-nbf-and-nbnm-individuals&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;At the start of the projection, we need the counts of NBF and NBNM individuals by age and province — a level of detail INE does not publish. To establish a coherent baseline, I draw on ECEPOV 2021-microdata and proceed as follows.&lt;&#x2F;p&gt;
&lt;p&gt;First, I classify every repsondent in the ECEPOV-2021 sample into one of three groups: (1) native-born individuals with a foreign-born mother (NBF), (2) native-born individuals with a native-born mother (NBNM), and (3) foreign-born individuals (FB). For all individuals born in Spain with unknown maternal birthplace, I follow a deliberatively conservative approach and classify them as having native-born mothers to avoid overestimating the size of the NBF group. Thus, the results presented here should be interpreted as a &lt;em&gt;lower bound&lt;&#x2F;em&gt; on the number of second-generation immigrants.&lt;&#x2F;p&gt;
&lt;p&gt;Now, remember that the goal is to produce population counts as of January 1 each year, but the fieldwork for ECEPOV was conducted in mid-2021. To simplify things, I treat the survey snapshot as if it were taken on January 1, 2022, and then project forward two full years of births and deaths to reach January 1, 2024 using the model described above. This generates preliminary age $\times$ province estimates for each subgroup of interest (NBF and NBNM), reflecting how populations would have evolved absent migration.&lt;&#x2F;p&gt;
&lt;p&gt;Second, because INE only reports the &lt;em&gt;total&lt;&#x2F;em&gt; population by age and province (including foreign‑born) — denoted $N^{INE}(p,a,2024)$ — I allocate that total across our three groups in the exact proportions implied by the ECEPOV projection. In practice:&lt;&#x2F;p&gt;
&lt;p&gt;$$NBF(p,a,2024) = \frac{\widehat NBF^{ECEPOV}(p,a,2024)}{\widehat N^{ECEPOV}(p,a,2024)} \times N^{INE}(p,a,2024)$$&lt;&#x2F;p&gt;
&lt;p&gt;and similarly for NBNM. Essentially, this rescaling preserves the relative proportions from the survey estimates.&lt;&#x2F;p&gt;
&lt;p&gt;Since the total number of NBF and NBNM individuals in each age $\times$ province cell at baseline are now defined, we have all the necessary data to project both populations forward. However, it is important to note that at every intermediate at every step of the projection — both in establishing the baseline and in year-to-year changes — the absolute numbers of NBF and NBNM individuals remain misaligned with INE&#x27;s official native-born population totals. This stems from two factors: (1) small differences between the native-born population shares in the ECEPOV survey and the true native-born shares in the population in 2021, and (2) the fact that the births-and-deaths projections model abstracts from migration flows. To correct for this, in the next step I re-anchor NBF and NBNM sums to INE’s native-born margins in every province and year using a scaling factor that absorbs any aggregate bias in absolute levels. Intuitively, what persists into the projection is the &lt;em&gt;within-native&lt;&#x2F;em&gt; age $\times$ province ratio of NBF to NBNM individuals, which becomes the true driver of the second-generation relative shares.&lt;&#x2F;p&gt;
&lt;!-- To calculate the relative population shares of NBF and NBNM individuals at baseline (2024), I simply apply the same algorithm to pre-baseline years (2021-2023) using the ECEPOV-2021 data, adding the historical numbers of births to foreign-born mothers. Let 

$$\widehat{NBF}^{ECEPOV}(p,a,2024)$$  

$$\widehat{NBNM}^{ECEPOV}(p,a,2024)$$ 

be the population of NBF and NBNM individuals in 2024 projected forward from the ECEPOV-2021 data, respectively. Then:

$$\widehat{NBF}(p,a,2024) = \frac{\widehat{NBF}^{ECEPOV}(p,a,2024)}{\widehat{N}^{ECEPOV}(p,a,2024)} \times N(p,a,2024)$$

$$\widehat{NBNM}(p,a,2024) = \frac{\widehat{NBNM}^{ECEPOV}(p,a,2024)}{\widehat{N}^{ECEPOV}(p,a,2024)} \times N(p,a,2024)$$

where $\widehat{N}^{ECEPOV}(p,a,2024)$ is the total population in 2024 projected forward from the ECEPOV-2021 data, and $N(p,a,2024)$ is the total population for each province and age group in 2024 from INE&#x27;s official projections. Since the total number of NBF and NBNM individuals in each age $\times$ province cell at the start of the projection period are now defined, we have all the necessary data to project both populations forward. --&gt;
&lt;h2 id=&quot;handling-migration&quot;&gt;Handling migration&lt;a class=&quot;zola-anchor&quot; href=&quot;#handling-migration&quot; aria-label=&quot;Anchor link for: handling-migration&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;We have established a basic births-and-deaths projection framework and calibrated the initial stock of NBF and NBNM individuals, which means we have all the necessary data to project both populations forward. But how do we incorporate migration? Rather than attempting to estimate $M_{NBF}(a,t)$ and $M_{NBNM}(a,t)$ directly, &lt;strong&gt;I simply assume that net migration inside the native-born block is allocated in proportion to pre-migration shares within each province.&lt;&#x2F;strong&gt; In other words, if the NBF population is 25% of the native-born population in a province before accounting for migration, I assume it will remain 25% after migration. To formalize this intuition mathematically, consider the following derivation.&lt;&#x2F;p&gt;
&lt;p&gt;Let&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;$\hat{N}(p,t)$ – birth-and-death projection for the total population (no migration)&lt;&#x2F;li&gt;
&lt;li&gt;$N^{INE}(p,t)$ – INE&#x27;s official projection for the total population (includes migration)&lt;&#x2F;li&gt;
&lt;li&gt;$\widehat{Native}(p,t)$ – birth-and-death projection for native-born population (no migration)&lt;&#x2F;li&gt;
&lt;li&gt;$Native^{INE}(p,t)$ – INE&#x27;s official projection for native-born population (includes migration)&lt;&#x2F;li&gt;
&lt;li&gt;$\widehat{NBF}(p,t)$ – birth-and-death projection for the NBF population&lt;&#x2F;li&gt;
&lt;li&gt;$NBF(p,t)$ – true, unknown NBF stock we want&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;Using INE&#x27;s projections for native and foreign-born populations, I calculate the net migration for the native-born population for each province and year:&lt;&#x2F;p&gt;
&lt;p&gt;$$M_{Native}(p,t) = Native^{INE}(p,t) - \widehat{Native}(p,t)$$&lt;&#x2F;p&gt;
&lt;p&gt;Within the native-born population, I assume that net migration is split in proportion to each subgroup&#x27;s share:&lt;&#x2F;p&gt;
&lt;p&gt;$$\frac{M_{NBF}(p,t)}{M_{Native}(p,t)} = \frac{\widehat{NBF}(p,t)}{\widehat{Native}(p,t)}$$&lt;&#x2F;p&gt;
&lt;p&gt;This can be rearranged to:&lt;&#x2F;p&gt;
&lt;p&gt;$$M_{NBF}(p,t) = M_{Native}(p,t) \times \frac{\widehat{NBF}(p,t)}{\widehat{Native}(p,t)}$$&lt;&#x2F;p&gt;
&lt;p&gt;Therefore:&lt;&#x2F;p&gt;
&lt;p&gt;$$NBF(p,t) = \widehat{NBF}(p,t) + M_{NBF}(p,t)$$
$$NBF(p,t) = \widehat{NBF}(p,t) \left(1 + \frac{M_{Native}(p,t)}{\widehat{Native}(p,t)}\right)$$
$$NBF(p,t) = \widehat{NBF}(p,t) \frac{Native^{INE}(p,t)}{\widehat{Native}(p,t)}$$&lt;&#x2F;p&gt;
&lt;p&gt;Define a province-specific scaling factor for natives:&lt;&#x2F;p&gt;
&lt;p&gt;$$\lambda_{Native}(p,t) = \frac{Native^{INE}(p,t)}{\widehat{Native}(p,t)}$$&lt;&#x2F;p&gt;
&lt;p&gt;To get the final NBF population, we can then simply multiply:&lt;&#x2F;p&gt;
&lt;p&gt;$$\boxed{NBF(p,t) = \widehat{NBF}(p,t) \times \lambda_{Native}(p,t)}$$&lt;&#x2F;p&gt;
&lt;p&gt;By anchoring in this way at each projection step, I eliminate any level distortions arising both from the initial two-year projection of ECEPOV-derived shares to the January 1 2024 baseline and from each subsequent year-to-year update, and ensure the NBF and NBNM totals match the native-born population from INE&#x27;s official projections.&lt;&#x2F;p&gt;
&lt;!-- ## Final adjustments

After calculating the natural population dynamics, I apply the migration adjustment by comparing our projected totals with INE&#x27;s official projections:

$$\lambda(p,t+1) = \frac{N(p,t+1)}{\sum_{a} \hat{N}(p,a,t+1)}$$

$$NBF_{\text{adjusted}}(a,t) = \sum_{a} \widehat{NBF}(p,a,t+1) \times \lambda(p,t+1)$$ --&gt;
&lt;!-- Finally, we calculate the share of native-born with foreign-born mothers in each province:

$$Share_{NBF}(p,t+1) = \frac{NBF_{\text{adjusted}}(a,t)}{N(p,t+1)}$$ --&gt;
&lt;h1 id=&quot;results&quot;&gt;Results&lt;a class=&quot;zola-anchor&quot; href=&quot;#results&quot; aria-label=&quot;Anchor link for: results&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;The results for the whole country are shown in &lt;a href=&quot;#fig1&quot;&gt;Figure 1&lt;&#x2F;a&gt;. The share of native-born individuals with native-born mothers will decrease from 77.1% in 2024 to just 62.3% by 2039. In just 15 years, the traditionally dominant demographic group will lose around 15 percentage points of its population. Meanwhile, first-generation immigrants will surge from 18.1% to 28.6% of the population, while second-generation immigrants (native-born with foreign-born mothers) will grow from 4.8% to 9.1%.&lt;&#x2F;p&gt;
&lt;p&gt;Put these numbers together, and you get a striking conclusion: &lt;strong&gt;by 2039, nearly 4 in 10 Spanish residents will be either immigrants themselves or the children of immigrants&lt;&#x2F;strong&gt;. This represents one of the most rapid demographic transformations in modern European history.&lt;&#x2F;p&gt;
&lt;figure id=&quot;fig1&quot;&gt;
  &lt;figcaption&gt;Figure 1 &lt;&#x2F;figcaption&gt;
  &lt;img src=&quot;&#x2F;img&#x2F;pop_composition_total_nacional.png&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; border: 0px none;&quot;&gt;&lt;&#x2F;iframe&gt;
&lt;&#x2F;figure&gt;
&lt;p&gt;However, the national picture masks a profound regional divergence. As &lt;a href=&quot;#fig2&quot;&gt;Figure 2&lt;&#x2F;a&gt; reveals, Spain will fracture into two distinct demographic realities.&lt;&#x2F;p&gt;
&lt;p&gt;In Southern provinces like Córdoba, Jaén, and Badajoz — areas with traditionally agrarian economies and fewer economic opportunities — the demographic composition will remain relatively stable. Here, native-born individuals with native-born mothers will still comprise close to 90% of the population by 2039, preserving much of the traditional demographic character. &lt;strong&gt;But along the Mediterranean coast and in economic powerhouses like Madrid and Barcelona, a dramatically different Spain is emerging: in these regions, immigrant and second-generation populations combined will approach or even exceed 50% of the total population by 2039&lt;&#x2F;strong&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;In Spain&#x27;s most economically vital and populous areas, the native-origin population will be close to becoming a minority in less than 15 years. Absent any major shifts in migration patterns, the municipalities of Madrid and Barcelona will transform into multicultural hubs more reminiscent of London or New York than the Spain of previous generations.&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#2&quot;&gt;2&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;&lt;&#x2F;p&gt;
&lt;!-- However, the patterns will vary significantly across provinces (&lt;a href=&quot;#fig2&quot;&gt;Figure 2&lt;&#x2F;a&gt;). Echoing Spanish history, it appears the country will be once again divided into two distinct realities, though this time along demographic (rather than ideological) lines. Provinces like Córdoba, Jaén or Badajoz will maintain their traditional demographic composition, with native-born individuals with native-born mothers still comprising around 80% of the population by 2039. Meanwhile, coastal and economic hubs like Madrid, Barcelona, Girona, and the Balearic Islands are headed toward a fundamentally different composition, where immigrant and second-generation populations will approach or exceed 50% of the total population. --&gt;
&lt;figure id=&quot;fig2&quot;&gt;
  &lt;figcaption&gt;Figure 2 &lt;&#x2F;figcaption&gt;
  &lt;img src=&quot;&#x2F;img&#x2F;native_share_dumbbell.png&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; border: 0px none;&quot;&gt;&lt;&#x2F;iframe&gt;
&lt;&#x2F;figure&gt;
&lt;p&gt;The scale of this shift becomes even more striking when we consider the age distribution. By 2039, the immigrant-origin population (first and second generation combined) will be substantially overrepresented among Spain&#x27;s working-age population, while the native-origin population will be disproportionately concentrated among older age groups.&lt;&#x2F;p&gt;
&lt;p&gt;We can do a simple back-of-the-envelope calculation to illustrate this point. In 2024, foreign-born individuals constituted 23.1% of Spain&#x27;s working-age population (ages 15-64), despite representing only 18.2% of the total population—an overrepresentation factor of 1.27. Projecting this pattern forward, by 2039, foreign-born individuals will comprise approximately 36.3% of Spain&#x27;s working-age population (28.6% $\times$ 1.27).&lt;&#x2F;p&gt;
&lt;p&gt;Among the remaining native-born segment (63.7% of the working-age population), second-generation immigrants will represent about 10.3%, translating to 6.6% of the total working-age population. &lt;strong&gt;When combined, these figures imply that, by 2039, approximately 43% of Spain&#x27;s workforce — over one in four working-age individuals — will be either first or second-generation immigrants.&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;h2 id=&quot;some-external-validation&quot;&gt;Some external validation&lt;a class=&quot;zola-anchor&quot; href=&quot;#some-external-validation&quot; aria-label=&quot;Anchor link for: some-external-validation&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;At this point, you might be wondering how well these projections align with other estimates available and whether they are reasonable at all.&lt;&#x2F;p&gt;
&lt;p&gt;While it is tricky to validate the long-term projections against other estimates since the latter are simply not available (or at least I am not aware of any of them), we can still cross-check the estimates at the 2024 baseline against historical data. In particular, since 2021, Eurostat has been publishing data on population totals for different segments of the working-age population by migration group (see &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;ec.europa.eu&#x2F;eurostat&#x2F;databrowser&#x2F;view&#x2F;lfsa_pgaccpm__custom_16608096&#x2F;default&#x2F;table?lang=en&amp;amp;page=time:2021&quot;&gt;here&lt;&#x2F;a&gt;). In 2024, Eurostat reports that native-born individuals with at least one foreign-born parent (they don’t break out figures for foreign-born mothers alone) constitute 3.80% of Spain&#x27;s working-age population. To be extremely tedious, I recalculated my projections using the same definition — either parent foreign-born rather than just mothers — and obtained 3.75% for the same year. This is a difference of only 0.05 percentage points relative to official counts, which is well within the margin of error.&lt;&#x2F;p&gt;
&lt;!-- If we compare the 2024 projection of working-age native-born individuals with foreign-born mothers (2.6%) with Eurostat&#x27;s figures for working-age native-born individuals with at least one foreign-born parent in that same year (3.80%), the alignment is quite reasonable given the definitional gap (remember that I only track foreign-born mothers rather than either parent being foreign-born -- hence the lower figure in the projections). To be extremely tedious, I checked the numbers in the projection if I had used either parent being foreign-born instead of just the mother, and I get 3.75% for 2024. This is a difference of only 0.05 percentage points relative to official counts, which is well within the margin of error. --&gt;
&lt;!-- That said, let me finally note that no forecasting exercise is perfect, and obsessing over the third decimal place risks giving a false sense of precision. I like to think about this type of analysis as a tool for framing plausible scenarios, not a contest to nail the exact numbers. If results are within a reasonable ballpark and transparently derived, they have done their job — providing the public with a quantitative foundation for discussing deeper questions that these simple mechanical exercises cannot answer on their own. --&gt;
&lt;h1 id=&quot;conclusion&quot;&gt;Conclusion&lt;a class=&quot;zola-anchor&quot; href=&quot;#conclusion&quot; aria-label=&quot;Anchor link for: conclusion&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;What we&#x27;re witnessing in Spain is nothing short of a demographic revolution. In just over 40 years (1998-2039), Spain will have transformed from a country where foreign-born residents were a tiny minority (2.9%) to one where nearly 40% of residents will be either immigrants themselves or their children.&lt;&#x2F;p&gt;
&lt;p&gt;One cannot help but wonder about the broader implications of such a dramatic shift -- the social reality these projections predict will inevitably reshape Spanish society. How will regional identities evolve when the composition of the population changes so fundamentally? What will happen, for instance, to regional languages like Catalan or Basque, which have historically been maintained through intergenerational transmission among native populations? And what happens when different regions within the same country experience such divergent demographic trajectories?&lt;&#x2F;p&gt;
&lt;p&gt;These are empirical questions that only time will answer, but history suggests that rapid demographic changes often precede significant social and political realignments. Spain, in this sense, may serve as an important case study for other European nations in the coming decades.&lt;&#x2F;p&gt;
&lt;hr &#x2F;&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;1&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;1&lt;&#x2F;sup&gt;
&lt;p&gt;If you are curious, I developed an interactive visualization tool to explore data on the international migrant stock in Spain since 1998 to the present using microdata from administrative municipal registers and censuses. Check it out: &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;http:&#x2F;&#x2F;dedonde.es&#x2F;&quot;&gt;http:&#x2F;&#x2F;dedonde.es&#x2F;&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;2&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;2&lt;&#x2F;sup&gt;
&lt;p&gt;With the caveat that, compared to London or New York, origin countries are much more concentrated in Spanish cities. For example, individuals born in Latin America accounted for roughly 48% of the total foreign-born population stock at the national level in 2024.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
</content>
	</entry>
	<entry xml:lang="en">
		<title>Montoya va donde brilla (in praise of not belonging)</title>
		<published>2025-02-09T00:00:00+00:00</published>
		<updated>2025-02-09T00:00:00+00:00</updated>
		<link href="https://pablogguz.github.io/blog/dondebrilla/" type="text/html"/>
		<id>https://pablogguz.github.io/blog/dondebrilla/</id>
		<content type="html">&lt;p&gt;Let’s start with the circus. &lt;em&gt;&lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;es.wikipedia.org&#x2F;wiki&#x2F;La_isla_de_las_tentaciones&quot;&gt;La Isla de las Tentaciones&lt;&#x2F;a&gt;&lt;&#x2F;em&gt; — Spain’s answer to &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Temptation_Island_(TV_series)&quot;&gt;Temptation Island&lt;&#x2F;a&gt; — is not a reality TV show. It’s a social experiment designed to strip love down to its most pathetic, voyeuristic core. The premise is simple: take a handful of couples, separate them from their partners on a tropical paradise, and throw in attractive singles to test their loyalty. The result? A spectacle of jealousy, cheating, and &lt;strong&gt;the kind of emotional carnage that makes you wonder why anyone would willingly sign up for this&lt;&#x2F;strong&gt;.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;who-is-montoya&quot;&gt;Who is Montoya?&lt;a class=&quot;zola-anchor&quot; href=&quot;#who-is-montoya&quot; aria-label=&quot;Anchor link for: who-is-montoya&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;&lt;em&gt;La Isla de las Tentaciones&lt;&#x2F;em&gt; operates on the fundamental principle that human relationships can be reduced to game theory. &lt;strong&gt;Put a bunch of young, conventionally hot people in isolation, add alcohol and cameras, and watch as they navigate their particular prisoner&#x27;s dilemma: cheat or be cheated on&lt;&#x2F;strong&gt;. The participants, selected for their genetic lottery wins and Instagram follower counts, perform the rituals of courtship under manufactured conditions. It might not be the most honest representation of contemporary romance, but something about it feels uncomfortably familiar.&lt;&#x2F;p&gt;
&lt;p&gt;In Season 8, &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.thesun.co.uk&#x2F;tvandshowbiz&#x2F;33254205&#x2F;temptation-island-spain-star-jose-carlos-montoya&#x2F;&quot;&gt;José Carlos Montoya&lt;&#x2F;a&gt; became the lab rat we deserved. His girlfriend, Anita, did what anyone would do under fluorescent lights and the gaze of a million strangers: &lt;strong&gt;she fucked someone else&lt;&#x2F;strong&gt;. Montoya’s response? &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;x.com&#x2F;PopCulture2000s&#x2F;status&#x2F;1886875821253476689&quot;&gt;He ran&lt;&#x2F;a&gt;. Not toward dignity, but into his own humiliation, slapping sand as producers barked &lt;em&gt;Montoya, por favor!&lt;&#x2F;em&gt; like they were trying to stop a drunk friend from doing something stupid.&lt;&#x2F;p&gt;
&lt;p&gt;The cameras loved it. We loved it. The clip went viral. Of course it did. The editing was exquisite – Spielberg himself could not have staged it better. The camera angles, the lighting, the rhythmic cuts between Montoya&#x27;s crumbling face and his girlfriend&#x27;s... alternative activities. &lt;strong&gt;Pure cinema&lt;&#x2F;strong&gt;.&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;figure id=&quot;fig1&quot;&gt;
  &lt;img src=&quot;&#x2F;img&#x2F;montoya.jpg&quot; loading=&quot;lazy&quot; style=&quot;width: 90%; border: 0px none;&quot;&gt;
&lt;&#x2F;figure&gt;
&lt;br&gt; 
&lt;p&gt;But before this pivotal moment, Montoya had already established himself as the show&#x27;s main character through a series of fantastic one-liners, performances, and a general air of tragicomedy. One of his most memorable moments was when, during a breakdown, he declared with the solemnity of a Shakespearean actor that &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;x.com&#x2F;GusmaoAne&#x2F;status&#x2F;1878932825854169307&quot;&gt;&lt;strong&gt;&lt;em&gt;&quot;Montoya va donde brilla&quot;&lt;&#x2F;em&gt;&lt;&#x2F;strong&gt;&lt;&#x2F;a&gt; (&quot;Montoya goes where he shines&quot;).&lt;&#x2F;p&gt;
&lt;p&gt;At first, it seemed like a linguistic flub — an unintended inversion of the phrase &lt;strong&gt;&lt;em&gt;&quot;Montoya brilla donde va&quot;&lt;&#x2F;em&gt;&lt;&#x2F;strong&gt; (&quot;Montoya shines wherever he goes&quot;, which is a common Spanish idiom to say that someone is able to thrive in any environment). But as the meme spread, I realized that this accidental catchphrase was never meant to be a joke. &lt;strong&gt;It was a manifesto&lt;&#x2F;strong&gt;.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;the-accidental-prophet&quot;&gt;The accidental prophet&lt;a class=&quot;zola-anchor&quot; href=&quot;#the-accidental-prophet&quot; aria-label=&quot;Anchor link for: the-accidental-prophet&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;&lt;em&gt;Shine wherever you go&lt;&#x2F;em&gt;. Society loves this mantra. It’s the self-help gospel of grinding, hustling, and proving your worth in hostile spaces. &lt;em&gt;Be brilliant, be visible, be universally admired&lt;&#x2F;em&gt;. The kind of self-optimization propaganda that keeps LinkedIn influencers employed and therapy offices full. But &lt;em&gt;&quot;va donde brilla&quot;&lt;&#x2F;em&gt; suggests a quieter rebellion: &lt;strong&gt;what if life isn’t about forcing yourself to shine in every room, but about finding the rooms that already see your light?&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;p&gt;There is a difference between adapting and contorting. While society pushes people to be everything to everyone, the truth is uglier: &lt;strong&gt;our talents, love, and vulnerabilities only matter when met with eyes soft enough to behold them&lt;&#x2F;strong&gt;. Some places will never be home. Some people will never see you. Some loves will never fit, no matter how much skin you scrape off trying to squeeze into them. &lt;strong&gt;We&#x27;ve all seen it — those relationships where one partner tolerates the other like an interesting furniture choice they&#x27;ve grown to regret.&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;p&gt;To &lt;em&gt;go where you shine&lt;&#x2F;em&gt; is to reject the tyranny of universal validation. In an era obsessed with optimizing ourselves into bland, marketable versions of humanity, Montoya is a man brave enough to say that &lt;em&gt;he is not for everyone, and that&#x27;s okay&lt;&#x2F;em&gt;. Your little corner of the universe might not look like success — his didn’t either. He stumbled upon it on a reality show designed to exploit his insecurities and a viral clip that turned his heartbreak into performance art. &lt;strong&gt;In an accidental margin of the internet, he found a space that let him be messy, and tragic, and beautiful, all at once&lt;&#x2F;strong&gt;. A space that didn’t ask him to be anything other than what he is: &lt;strong&gt;completely human&lt;&#x2F;strong&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;The rooms that deserve you are out there. They&#x27;re softly lit by old lamps, smell like Sunday morning coffee, and most likely the chairs don&#x27;t match. You’ll know them by the way your heart sings when you walk in.&lt;&#x2F;p&gt;
&lt;p&gt;And if it doesn’t, well.&lt;&#x2F;p&gt;
&lt;p&gt;You can always run.&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;figure id=&quot;fig1&quot;&gt;
  &lt;img src=&quot;&#x2F;img&#x2F;montoyagif.gif&quot; loading=&quot;lazy&quot; style=&quot;width: 90%; border: 0px none;&quot;&gt;
&lt;&#x2F;figure&gt;
&lt;br&gt; 
&lt;br&gt;
&lt;!-- Montoya found his light in the most unlikely place — a reality show designed to exploit his insecurities and a viral clip that turned his heartbreak into performance art. In a dim corner of the internet, he found the space that let him be pathetic, and tragic, and beautiful, all at once. The space that didn’t ask him to be anything other than what he is: raw, messy, and completely human. --&gt;
&lt;!-- Your little corner of the universe might not look like success. It might look like failure, like mess, like a reality show clip shared a million times. But in that corner, at least you get to be gloriously, catastrophically yourself. --&gt;
&lt;!-- &lt;blockquote class=&quot;callout tip&quot;&gt;
    
    &lt;div class=&quot;icon&quot;&gt;
        &lt;svg xmlns=&quot;http:&#x2F;&#x2F;www.w3.org&#x2F;2000&#x2F;svg&quot; viewBox=&quot;0 0 24 24&quot; width=&quot;20&quot; height=&quot;20&quot;&gt;&lt;path d=&quot;M9.97308 18H11V13H13V18H14.0269C14.1589 16.7984 14.7721 15.8065 15.7676 14.7226C15.8797 14.6006 16.5988 13.8564 16.6841 13.7501C17.5318 12.6931 18 11.385 18 10C18 6.68629 15.3137 4 12 4C8.68629 4 6 6.68629 6 10C6 11.3843 6.46774 12.6917 7.31462 13.7484C7.40004 13.855 8.12081 14.6012 8.23154 14.7218C9.22766 15.8064 9.84103 16.7984 9.97308 18ZM10 20V21H14V20H10ZM5.75395 14.9992C4.65645 13.6297 4 11.8915 4 10C4 5.58172 7.58172 2 12 2C16.4183 2 20 5.58172 20 10C20 11.8925 19.3428 13.6315 18.2443 15.0014C17.624 15.7748 16 17 16 18.5V21C16 22.1046 15.1046 23 14 23H10C8.89543 23 8 22.1046 8 21V18.5C8 17 6.37458 15.7736 5.75395 14.9992Z&quot; fill=&quot;currentColor&quot;&gt;&lt;&#x2F;path&gt;&lt;&#x2F;svg&gt;
    &lt;&#x2F;div&gt;
    &lt;div class=&quot;content&quot;&gt;
        
        &lt;p&gt;&lt;strong&gt;Watch the full clip&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
        
        &lt;p&gt;You can watch Montoya&#x27;s iconic moment &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;x.com&#x2F;PopCulture2000s&#x2F;status&#x2F;1886875821253476689&quot;&gt;here&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;

    &lt;&#x2F;div&gt;
&lt;&#x2F;blockquote&gt; --&gt;
&lt;hr &#x2F;&gt;
&lt;!-- [^1] This is not, by any means, a call to complacency. Growth is painful, and sometimes it requires us to be in uncomfortable situations. But there’s a difference between being uncomfortable because you’re growing and being uncomfortable because you’re trying to fit into a space that was never meant for you. It also goes without saying that this is not an excuse to justify toxic behavior. If you are a terrible person, you should definitely work on that.

[^2] This isn&#x27;t an invitation to passive contemplation. Nothing that is worth having will fall into your lap! You should definitely put yourself out there, try new things, meet new people, learn obsessively. But do it from a place of genuine curiosity rather than plain validation-seeking: when you chase experiences to please others or fill a void, you&#x27;re still hunting. Your garden won&#x27;t flourish if you become a hermit and never leave the house, but it also won&#x27;t thrive if you&#x27;re only planting what you think others want to see. --&gt;
&lt;!-- Don&#x27;t get me wrong. Within reasonable constraints, we _must_ be adaptable and resilient. We have to learn navigate different environment and social codes, and we all should strive to be the best version of ourselves in every possible realm. But there’s a difference between adapting and contorting. While society somehow pushes people to be everything to everyone, the truth is uglier: at the margin, our talents, love, and vulnerabilities only matter when met with eyes soft enough to behold them. Some places will never be home. Some people will never see you. Some loves will never fit, no matter how much skin you scrape off trying to squeeze into them. You can spend years trying to shine in an office that finds you mildly irritating, or in a relationship where your partner tolerates you like an interesting furniture choice they&#x27;ve grown to regret. But trying to fit where you are not supposed to does not lead to personal growth. It leads to self-erasure.[^1] --&gt;
&lt;!-- People often frame relationships as a hunt. _Get that partner. Get that job. Get that life_. This mindset reduces connection to a game of conquest. However, hunting with the sole purpose of being liked is not where true connections are found: it’s where you lose yourself in the chase. Instead, successful human interactions resemble less a hunt and more the slow, obsessive cultivation of a peculiar garden. Genuine connection isn&#x27;t found in a desperate sprint toward universal appeal, but it emerges in the patient cultivation of your own ecosystem. Your interests, your quirks, your specific way of seeing the world — these are rare species to be nurtured. And when you start flourishing, the right people find their way to your garden.[^2] --&gt;
&lt;!-- Find your little corner of the universe where your flaws are not just tolerated but revered. Let your love be small. Let it be specific. Let it be desperate. Let it be yours, and of no one else. --&gt;</content>
	</entry>
	<entry xml:lang="en">
		<title>Wellbeing I: Introduction</title>
		<published>2024-03-22T00:00:00+00:00</published>
		<updated>2024-03-22T00:00:00+00:00</updated>
		<link href="https://pablogguz.github.io/blog/wellbeing-i-intro/" type="text/html"/>
		<id>https://pablogguz.github.io/blog/wellbeing-i-intro/</id>
		<content type="html">&lt;img src=&quot;&#x2F;img&#x2F;wellbeing_i_picv3.webp&quot; width=&quot;500&quot;&#x2F;&gt;
&lt;br&gt;
&lt;blockquote class=&quot;quote&quot;&gt;
    
    &lt;div class=&quot;icon&quot; style=&quot;display: none;&quot;&gt;&lt;svg fill=&quot;currentColor&quot; xmlns=&quot;http:&#x2F;&#x2F;www.w3.org&#x2F;2000&#x2F;svg&quot;  width=&quot;12&quot; height=&quot;12&quot; viewBox=&quot;796 698 200 200&quot;&gt;
&lt;g&gt;
	&lt;path d=&quot;M885.208,749.739v-40.948C836.019,708.791,796,748.81,796,798v89.209h89.208V798h-48.26
		C836.948,771.39,858.598,749.739,885.208,749.739z&quot;&#x2F;&gt;
	&lt;path d=&quot;M996,749.739v-40.948c-49.19,0-89.209,40.019-89.209,89.209v89.209H996V798h-48.26
		C947.74,771.39,969.39,749.739,996,749.739z&quot;&#x2F;&gt;
&lt;&#x2F;g&gt;
&lt;&#x2F;svg&gt;&lt;&#x2F;div&gt;
    &lt;div class=&quot;content&quot;&gt;&lt;p&gt;&lt;em&gt;GDP measures everything, in short, except that which makes life worthwile&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
    
    &lt;div class=&quot;from&quot;&gt;
      &lt;p&gt;— Robert F. Kennedy&lt;&#x2F;p&gt;

    &lt;&#x2F;div&gt;
    
  &lt;&#x2F;blockquote&gt;
&lt;p&gt;A brilliant surgeon has five patients, each in need of a different organ transplant. All five will die without these transplants, and suitable organs are not available. In the next room, there is a healthy person who came in for a routine checkup. The surgeon knows they could use this person’s organs to save the five patients. No one would ever suspect the cause of the healthy person&#x27;s demise.&lt;&#x2F;p&gt;
&lt;p&gt;What should the surgeon do? Should they kill the healthy person to save the five patients?&lt;&#x2F;p&gt;
&lt;p&gt;This is a variation of the famous &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Trolley_problem&quot;&gt;trolley problem&lt;&#x2F;a&gt;, a thought experiment in ethics and moral psychology. The problem is designed to highlight the tension between two moral principles: the principle of &lt;strong&gt;utilitarianism&lt;&#x2F;strong&gt;, which states that the right action is the one that maximizes overall wellbeing, and the principle of &lt;strong&gt;deontology&lt;&#x2F;strong&gt;, which states that the right action is the one that respects individual rights and duties. In the trolley problem, the utilitarian answer is to kill the healthy person to save the five patients, while the deontological answer is to not kill the healthy person, as doing so would violate their right to life.&lt;&#x2F;p&gt;
&lt;p&gt;I am not a utilitarian. However, philosophies of justice (like liberalism) often become unhelpful in scenarios where wellbeing trade-offs are inevitable. For instance, consider now that the surgeon has healthy organs available, but only enough to save three of the five patients. What should the surgeon do in that case?&lt;&#x2F;p&gt;
&lt;p&gt;In the real world, policymakers often face decisions that entail such compromises in one way or another: for example, should we invest in a new hospital, or in a new school? Should a government implement regulations to protect consumers, which could increase the cost of goods and services and potentially stifle innovation? Should a city invest in improving public transportation, potentially reducing traffic and pollution, but at the cost of higher taxes? These are all questions that involve trade-offs between different individuals&#x27; wellbeing, and more often than not, they are not easy to answer.&lt;&#x2F;p&gt;
&lt;p&gt;One of the ultimate goals of policy should be to maximize wellbeing &lt;em&gt;within constraints&lt;&#x2F;em&gt; for the individuals that it serves to. Philosophies of justice can help us delimit those constraints, but they are not enough to guide us on how to identify wellbeing maximizing interventions and balance wellbeing trade-offs.&lt;&#x2F;p&gt;
&lt;p&gt;And this is where wellbeing research comes in.&lt;&#x2F;p&gt;
&lt;p&gt;This is the first post of a series about wellbeing measurement and policy.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;definitions&quot;&gt;Definitions&lt;a class=&quot;zola-anchor&quot; href=&quot;#definitions&quot; aria-label=&quot;Anchor link for: definitions&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;Let&#x27;s start with the basics. What is wellbeing?&lt;&#x2F;p&gt;
&lt;p&gt;There is no universal definition. Wellbeing is a complex, multifaceted concept that encompasses a variety of domains, including emotional, psychological, and social aspects of life. It is also a subjective concept, meaning that it is defined by the individual&#x27;s own perception of their life, rather than by an external observer. This is important to keep in mind, as it means that wellbeing is not something that can be imposed from the outside, but rather something that needs to be elicited from the individual themselves.&lt;&#x2F;p&gt;
&lt;p&gt;In the context of policy research, I like to define wellbeing as the individual&#x27;s &lt;em&gt;inner subjective state of satisfaction with their own existence&lt;&#x2F;em&gt;. Nothing less, nothing more.&lt;&#x2F;p&gt;
&lt;p&gt;People often interchangeably use the idea of &lt;em&gt;happiness&lt;&#x2F;em&gt; to refer to wellbeing. I do not find this entirely accurate: happiness is a transient, momentary state of mind, while wellbeing is a more stable subjective state that factors in long-term mental and physical health, social connections, purpose, and fulfillment. However, happiness and other common wellbeing measures (such as life satisfaction) are typically highly correlated in the data, and the distinction is often blurred in practice.&lt;&#x2F;p&gt;
&lt;p&gt;One of the most common ways to elicit individual wellbeing is via large-scale household surveys. These surveys typically ask respondents to rate their own wellbeing, while at the same time collect a wide range of information about the respondent&#x27;s life circumstances, such as income, employment history, health, and social relationships. Such rich data allow us to explore what are the determinants of wellbeing at the individual level, and help us to understand better how different policies and interventions would affect people&#x27;s lives.&lt;&#x2F;p&gt;
&lt;p&gt;Broadly speaking, wellbeing measures can be split into three categories:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Evaluative&lt;&#x2F;strong&gt; measures stem from asking the respondent how satisfied or happy they are with their life. A common framing is &lt;em&gt;All thing considered, how satisfied&#x2F;happy are you with your life nowadays?&lt;&#x2F;em&gt;, with the answer options being either a continuous scale from 0-10 or a Likert-type scale.&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#1&quot;&gt;1&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt; A variant of the former approach is to ask respondents to rate their life on a ladder, where the top of the ladder represents the best possible life for them, and the bottom of the ladder represents the worst possible life for them. This is commonly known as the &lt;em&gt;Cantril ladder&lt;&#x2F;em&gt;, which is used in many international household surveys such as the &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.gallup.com&#x2F;analytics&#x2F;318875&#x2F;global-research.aspx&quot;&gt;Gallup World Poll&lt;&#x2F;a&gt;.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Hedonic&lt;&#x2F;strong&gt; approaches focus on measuring how people feel at a particular point in time. This is typically done by asking respondents to report their feelings at the moment of the survey, or to recall their feelings over a specific period of time. Experienced happiness is the most common measure of hedonic wellbeing, but negative emotions such as sadness, anger, and stress are also used.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Eudaimonic&lt;&#x2F;strong&gt; measures are based on the Aristotelian tradition with the idea that wellbeing is not just about feeling good, but also about functioning well. This approach focuses on the extent to which people are able to realize their potential, and to live a life that is meaningful and fulfilling. This is typically measured by asking respondents about their sense of purpose, their autonomy, and their relationships with others.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;Tipically, evaluative and hedonic measures are the most widely used in policy research. Eudaimonic measures, although philosophically appealing, are harder to interpret – for example, when we talk about being virtuous, it is not clear whether that should be an outcome in itself or another determinant (among many) of wellbeing.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;a href=&quot;#fig1&quot;&gt;Figure 1&lt;&#x2F;a&gt; shows cross-country averages of life satisfaction, as measured by the Cantril ladder question from the Gallup World Poll. The figure shows that there is a wide variation in life satisfaction across countries, with the highest levels of life satisfaction being found in the Nordic countries, and the lowest levels of life satisfaction being found in sub-Saharan Africa and the Middle East. The difference between the country with the highest average life satisfaction (Finland) and the country with the lowest average life satisfaction (Afghanistan) is about 6 points on the 0-10 Cantril ladder scale.&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;figure id=&quot;fig1&quot;&gt;
  &lt;figcaption&gt;Figure 1: Life satisfaction across countries &lt;&#x2F;figcaption&gt;
  &lt;iframe src=&quot;https:&#x2F;&#x2F;ourworldindata.org&#x2F;grapher&#x2F;happiness-cantril-ladder&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; height: 600px; border: 0px none;&quot;&gt;&lt;&#x2F;iframe&gt;
&lt;&#x2F;figure&gt;
&lt;br&gt;
&lt;h2 id=&quot;it-is-not-all-about-the-money&quot;&gt;It is (not) all about the money&lt;a class=&quot;zola-anchor&quot; href=&quot;#it-is-not-all-about-the-money&quot; aria-label=&quot;Anchor link for: it-is-not-all-about-the-money&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;Traditionally, income has been used as the main proxy for wellbeing in academic and policy circles. This is not surprising: income is relatively easy to measure, and we know that it is also good proxy for many things that directly affect wellbeing, such as &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;ourworldindata.org&#x2F;grapher&#x2F;life-expectancy-un-vs-gdp-per-capita-wb&quot;&gt;higher life expectancy&lt;&#x2F;a&gt;, better &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;ourworldindata.org&#x2F;grapher&#x2F;learning-outcomes-vs-gdp-per-capita&quot;&gt;education systems&lt;&#x2F;a&gt;, well-functioning &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;ourworldindata.org&#x2F;grapher&#x2F;labor-productivity-vs-gdp-per-capita&quot;&gt;labor markets&lt;&#x2F;a&gt;, political stability, absence of violence...&lt;&#x2F;p&gt;
&lt;p&gt;...and &lt;strong&gt;wellbeing itself&lt;&#x2F;strong&gt;. Across countries, higher GDP per capita is associated with greater levels of life satisfaction. This finding is remarkably consistent across surveys, and it holds up irrespective of the definition of wellbeing we use. &lt;a href=&quot;#fig2&quot;&gt;Figure 2&lt;&#x2F;a&gt; shows the relationship between GDP per capita and life satisfaction for a large number of countries, using data from the Gallup World Poll. High-income countries tend to have higher levels of life satisfaction than low-income countries, and the relationship is roughly log-linear. The latter is consistent with the idea of diminishing marginal utility of income: an extra dollar has more impact on wellbeing for the average Joe than for Jeff Bezos. Intuitive, right?&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;figure id=&quot;fig2&quot;&gt;
  &lt;figcaption&gt;Figure 2: Self-reported life satisfaction vs. GDP per capita &lt;&#x2F;figcaption&gt;
  &lt;iframe src=&quot;https:&#x2F;&#x2F;ourworldindata.org&#x2F;grapher&#x2F;gdp-vs-happiness&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; height: 600px; border: 0px none;&quot;&gt;&lt;&#x2F;iframe&gt;
&lt;&#x2F;figure&gt;
&lt;br&gt;
&lt;p&gt;What is the catch then? Why not just stick to income as a proxy for wellbeing?&lt;&#x2F;p&gt;
&lt;h3 id=&quot;the-easterlin-paradox&quot;&gt;The Easterlin paradox&lt;a class=&quot;zola-anchor&quot; href=&quot;#the-easterlin-paradox&quot; aria-label=&quot;Anchor link for: the-easterlin-paradox&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h3&gt;
&lt;p&gt;In 1974, economist Richard Easterlin established an empirical stylized fact:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;At a given point in time, richer people tend to be happier than poorer people, both within and across countries&lt;&#x2F;li&gt;
&lt;li&gt;However, as income grows and countries become richer, the average level of life satisfaction in such countries does not necessarily increase&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;This is known as the &lt;strong&gt;Easterlin paradox&lt;&#x2F;strong&gt; and has been a subject of much debate in the literature. Whereas (1) is undisputedly true, (2) is more ambiguous: whereas some countries have not experienced increases in average life satisfaction despite large increases in income, others have become happier as they have climbed the income ladder.&lt;&#x2F;p&gt;
&lt;p&gt;To illustrate the Easterlin paradox, consider &lt;a href=&quot;#fig3&quot;&gt;Figure 3&lt;&#x2F;a&gt;, which shows the evolution of GDP per capita and average life satisfaction (derived from large-scale longitudinal household surveys) for a sample of advanced economies. The figure shows that, despite considerable increases in within-country per capita income, the average level of life satisfaction has remained surprisingly constant over time. In the U.S., for example, GDP per capita has increased by more than 500% since 1950, but people are not any happier now than they were back then.&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;figure id=&quot;fig3&quot;&gt;
  &lt;figcaption&gt;Figure 3: The Easterlin Paradox&lt;&#x2F;figcaption&gt;
  &lt;img src=&quot;&#x2F;img&#x2F;clark_2018_paradox.png&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; height: 500px; border: 0px none;&quot;&gt;&lt;&#x2F;iframe&gt;
&lt;&#x2F;figure&gt;
&lt;p class=&quot;figure-note&quot;&gt;Source: &lt;a href=&quot;https:&#x2F;&#x2F;www.jstor.org&#x2F;stable&#x2F;j.ctvd58t1t&quot;&gt;Clark et al. (2018).&lt;&#x2F;a&gt;&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;p&gt;In light of the above, the answer to the question on why should we not rely solely on income as a proxy for wellbeing is straightforward: because it is simply not a good standalone proxy in the long-run.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;the-social-comparison-hypothesis&quot;&gt;The social comparison hypothesis&lt;a class=&quot;zola-anchor&quot; href=&quot;#the-social-comparison-hypothesis&quot; aria-label=&quot;Anchor link for: the-social-comparison-hypothesis&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h3&gt;
&lt;p&gt;One of the first explanations for the paradox was that income affects wellbeing through relative, rather than absolute, terms. In other words, what matters for determining wellbeing is how one&#x27;s income compares to that of others in society. This is known as the &lt;strong&gt;social comparison&lt;&#x2F;strong&gt; hypothesis and suggests that, if everyone becomes richer by a common factor, the average level of life satisfaction would not necessarily increase as the individuals&#x27; relative position within the income distribution would remain unchanged.&lt;&#x2F;p&gt;
&lt;p&gt;This is a very intuitive idea. If you are the only one in your social circle who has a Mercedes, you will probably feel very happy about it. But if all of a sudden all your friends start driving Ferraris, your Mercedes will now feel like a piece of junk. This is because the Mercedes serves as a status symbol: it is valuable insofar as it signals that you are in a better social position than your peers, but the moment it stops being exclusive it no longer serves that purpose.&lt;&#x2F;p&gt;
&lt;p&gt;There are a couple of interesting papers that provide evidence for the social comparison hypothesis. &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;http:&#x2F;&#x2F;www.jstor.org&#x2F;stable&#x2F;41724678&quot;&gt;Card et al. (2012)&lt;&#x2F;a&gt; run an experiment in which they randomly inform employees of the University of California about the salaries of their peers, and find that workers whose wage is below the median for their comparison group report lower job satisfaction and are more likely to start looking for a new job. &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.aeaweb.org&#x2F;articles?id=10.1257&#x2F;aer.20180819&quot;&gt;Perez-Truglia (2020)&lt;&#x2F;a&gt; exploits a natural experiment whereby Norwegian tax records became publicly available for everyone to see, and shows that income transparency increases the gap in happiness between high- and low-income individuals by 29 percent. This occurs because people at the bottom (top) of the income distribution typically overestimate (underestimate) their relative position, and increased transparency makes them realize how much worse (better) off they are compared to others.&lt;&#x2F;p&gt;
&lt;p&gt;In a less clean way, one could test for the social comparison hypothesis using a very simple regression framework in which we regress wellbeing on both individual income and the average income of a comparison group:&lt;&#x2F;p&gt;
&lt;p&gt;$$Y_i = \alpha + \beta X_i + \gamma \bar{X} + \epsilon_i$$&lt;&#x2F;p&gt;
&lt;p&gt;where $Y_i$ is individual $i$&#x27;s wellbeing, $X_i$ is individual $i$&#x27;s income, $\bar{X}$ is the average income of a group of comparators, and $\epsilon_i$ is the error term. If the relative income hypothesis holds, we would expect $\gamma$ to be negative, as individual wellbeing would decrease as the average income of the group of comparators increases.&lt;&#x2F;p&gt;
&lt;p&gt;This is exactly what we find. &lt;a href=&quot;#table1&quot;&gt;Table 1&lt;&#x2F;a&gt; shows the results of such a regression for cross-sections of different countries, in which the group of comparators is defined by people of the same sex, age group, region, and survey year, and the outcome is life satisfaction (measured on a 0-10 scale). Individual wellbeing is positively associated with individual income, but negatively associated with the average income of the comparison group. Except for the United States, a 1% increase in the average income of comparators reduces wellbeing in the same amount or more than a 1% increase in individual income increases wellbeing.&lt;&#x2F;p&gt;
&lt;div id=&quot;table1&quot;&gt;
&lt;figcaption&gt;Table 1: The social comparison hypothesis&lt;&#x2F;figcaption&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;&#x2F;th&gt;&lt;th&gt;Britain&lt;&#x2F;th&gt;&lt;th&gt;Germany&lt;&#x2F;th&gt;&lt;th&gt;Australia&lt;&#x2F;th&gt;&lt;th&gt;United States&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;Own income&lt;&#x2F;td&gt;&lt;td&gt;0.16 (.01)&lt;&#x2F;td&gt;&lt;td&gt;0.26 (.01)&lt;&#x2F;td&gt;&lt;td&gt;0.16 (.01)&lt;&#x2F;td&gt;&lt;td&gt;0.31 (.01)&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Comparator income&lt;&#x2F;td&gt;&lt;td&gt;-0.23 (.07)&lt;&#x2F;td&gt;&lt;td&gt;-0.25 (.04)&lt;&#x2F;td&gt;&lt;td&gt;-0.17 (.06)&lt;&#x2F;td&gt;&lt;td&gt;-0.19 (.03)&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;p class=&quot;figure-note&quot;&gt;Source: &lt;a href=&quot;https:&#x2F;&#x2F;www.cambridge.org&#x2F;gb&#x2F;universitypress&#x2F;subjects&#x2F;economics&#x2F;microeconomics&#x2F;wellbeing-science-and-policy&quot;&gt;Layard and De Neve (2023)&lt;&#x2F;a&gt;, Table 13.4. Standard errors in parentheses.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;br&gt;
&lt;p&gt;There are other potential explanations for the Easterlin paradox, such as the &lt;strong&gt;hedonic adaptation hypothesis&lt;&#x2F;strong&gt;. Roughly speaking, the idea here is that people adapt to their new income levels and their expectations increase (or decrease) accordingly. We will not discuss it in detail here, but there is equally interesting evidence on the role of adaptation in explaining the paradox.&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#2&quot;&gt;2&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;&lt;&#x2F;p&gt;
&lt;h3 id=&quot;effect-sizes&quot;&gt;Effect sizes&lt;a class=&quot;zola-anchor&quot; href=&quot;#effect-sizes&quot; aria-label=&quot;Anchor link for: effect-sizes&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h3&gt;
&lt;p&gt;Fine, so income growth does not necessarily lead to increases in average life satisfaction over time. But richer individuals are still happier than poorer individuals within a given cross-section. How big are the effects of income on wellbeing then? Are the people at the top of the income distribution much happier than those at the bottom?&lt;&#x2F;p&gt;
&lt;p&gt;Let&#x27;s start with some raw figures. In a simple regression framework whereby we regress wellbeing on income and a set of control variables, the partial correlation coefficients would tell us that income explains &lt;strong&gt;about 3%&lt;&#x2F;strong&gt; of the variation in wellbeing in advanced economies. This is not zero, but it is ridiculously low for it to be the most important determinant of wellbeing (as some people have traditionally argued). It is an important factor, but certainly not the only one.&lt;&#x2F;p&gt;
&lt;p&gt;For the sake of rigor, we should look at studies with quasi-experimental designs, which exploit exogenous shocks to income (e.g., lottery prizes) to estimate the causal effect of income on wellbeing. &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;doi.org&#x2F;10.1093&#x2F;restud&#x2F;rdaa006&quot;&gt;Lindqvist et al. (2020)&lt;&#x2F;a&gt; look at the long-run effects of lottery wealth on psychological wellbeing in Sweden, and find that an after-tax lottery prize of $100,000 increases overall life satisfaction by 0.037 standard deviations relative to the matched control group. To put it into perspective, that is about 70% of the difference in average life satisfaction between Denmark (the second happiest country in the world) and Iceland (the third happiest country in the world). In other words, this is a very small effect. Relatedly, the authors also explore the impact of lottery wealth on other outcomes, such as happiness and mental health, and in these cases the effects are not significantly different from zero. That is, lottery winners are not happier or less mentally distressed relative to the control group in the long-run.&lt;&#x2F;p&gt;
&lt;p&gt;As we saw in &lt;a href=&quot;#fig2&quot;&gt;Figure 2&lt;&#x2F;a&gt;, the effects of income on wellbeing are also non-linear. You probably have come across media articles arguing that, beyond a certain income threshold, the marginal effect of income on wellbeing is virtually zero. This is pretty much based on the (now sort of old) paper by &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;doi.org&#x2F;10.1073&#x2F;pnas.1011492107&quot;&gt;Kahneman and Deaton (2010)&lt;&#x2F;a&gt;, who use data for the U.S. and find that the relationship between income and &lt;em&gt;emotional wellbeing&lt;&#x2F;em&gt; is log-linear up to an annual income of $75,000, but flattens out beyond that point. The idea is that further increases in income do not necessarily improve the individual&#x27;s ability to engage in activities that are most important for their emotional wellbeing, such as spending time with friends and family, or having better health.&lt;&#x2F;p&gt;
&lt;p&gt;However, a key thing to note is that here we are looking at &lt;em&gt;hedonic&lt;&#x2F;em&gt; measures of wellbeing. Specifically, they use the average of three dicotomous, binary measures of positive affect (happiness, enjoyment, and frequent smiling and laughter) and the average of two binary measures of negative affect or &quot;blue&quot; states (worry and sadness). In contrast, the authors find that the relationship between income and &lt;em&gt;evaluative&lt;&#x2F;em&gt; wellbeing (i.e., what we have been referring to as life satisfaction or general happiness) shows no sign of satiation, at least to an amount over $160,000 per year.&lt;&#x2F;p&gt;
&lt;p&gt;Things got interesting when &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;doi.org&#x2F;10.1073&#x2F;pnas.2016976118&quot;&gt;Killingsworth (2021)&lt;&#x2F;a&gt; revisited this question and found that, conversely, hedonic wellbeing &lt;em&gt;does&lt;&#x2F;em&gt; rise with income well above the $75,000 per year threshold. The main argument of this paper is that the discrepancy with the original results might be explained by differences in the variable definitions: on top of just using binary variables (which are able to capture less variation), &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;doi.org&#x2F;10.1073&#x2F;pnas.1011492107&quot;&gt;Kahneman and Deaton (2010)&lt;&#x2F;a&gt; results are based on retrospective survey responses, which are subject to recall bias. Instead, &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;doi.org&#x2F;10.1073&#x2F;pnas.2016976118&quot;&gt;Killingsworth (2021)&lt;&#x2F;a&gt; uses large-scale real-time data collected through a smartphone app and a continuous scale to measure experienced wellbeing, finding that the latter increases linearly with log-income with no satiation point.&lt;&#x2F;p&gt;
&lt;p&gt;In a recent paper, the authors joined forces and reconciled the two previous findings (&lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;doi.org&#x2F;10.1073&#x2F;pnas.2208661120&quot;&gt;Killingsworth et al., 2023&lt;&#x2F;a&gt;). Here, they show that the flattening pattern in the relationship between income and experienced wellbeing depicted in the original study by Kahneman and Deaton &lt;em&gt;does exist&lt;&#x2F;em&gt;, but it only applies to a limited subset of the population. Using quantile regressions, they show that increases in income do not shift upwards the bottom 15-20% of the emotional wellbeing distribution (i.e., the least happy), but they do so across all other quantiles.&lt;&#x2F;p&gt;
&lt;p&gt;To illustrate this, see &lt;a href=&quot;#fig4&quot;&gt;Figure 4&lt;&#x2F;a&gt;. The happiness-income gradient is positive and significant for the top 80% of the distribution, but the slope turns flat for the those below the 15th percentile after $120,000 per year.&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;figure id=&quot;fig4&quot;&gt;
  &lt;figcaption&gt;Figure 4: Emotional wellbeing across the happiness distribution &lt;&#x2F;figcaption&gt;
  &lt;img src=&quot;&#x2F;img&#x2F;happy_quantiles.jpg&quot; loading=&quot;lazy&quot; style=&quot;width: 70%; border: 0px none;&quot;&gt;&lt;&#x2F;iframe&gt;
&lt;&#x2F;figure&gt;
&lt;p class=&quot;figure-note&quot;&gt;Source: &lt;a href=&quot;https:&#x2F;&#x2F;www.pnas.org&#x2F;doi&#x2F;epdf&#x2F;10.1073&#x2F;pnas.2208661120&quot;&gt;Killingsworth et al. (2023).&lt;&#x2F;a&gt;&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;p&gt;So, why &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;doi.org&#x2F;10.1073&#x2F;pnas.1011492107&quot;&gt;Kahneman and Deaton (2010)&lt;&#x2F;a&gt; claimed that the satiation pattern applies to the entire population? As mentioned above, the answer is in the variable definition. Since they used dichotomous measures, the authors ended up with a very skewed distribution of happiness, especially among higher incomes, suggesting a &quot;ceiling effect&quot; whereby the measure could not capture differences between higher degrees of happiness. In fact, the average reported positive affect for the range of high incomes in &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;doi.org&#x2F;10.1073&#x2F;pnas.1011492107&quot;&gt;Kahneman and Deaton (2010)&lt;&#x2F;a&gt; is 89% of a perfect score, but we know that the happy people are not all equally happy!&lt;&#x2F;p&gt;
&lt;p&gt;To see this, consider the example of a pass-or-fail math test, consisting of questions that most students can answer correctly. This test could be informative to identify the students who are doing badly, but it would not be able to differentiate between the students who are doing well because most of them would get the same pass grade. In a sense, the measures of affect in &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;doi.org&#x2F;10.1073&#x2F;pnas.1011492107&quot;&gt;Kahneman and Deaton (2010)&lt;&#x2F;a&gt; are just like this test: they are good at identifying the people who are unhappy, but they have limited ability to discriminate among levels of happiness. Instead of &lt;em&gt;&quot;Happiness increases with income, but it does not further increase above $75,000 per year&quot;&lt;&#x2F;em&gt;, it follows that the statement should have read &lt;em&gt;&quot;Unhappiness decreases with income, but it does not further decrease above $75,000 per year&quot;&lt;&#x2F;em&gt;.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;what-s-next&quot;&gt;What&#x27;s next?&lt;a class=&quot;zola-anchor&quot; href=&quot;#what-s-next&quot; aria-label=&quot;Anchor link for: what-s-next&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;In this post, I have briefly introduced the concept of wellbeing in the context of policy research, and disucssed the role of income in determining wellbeing. We have seen that, while income is an important driver of overall life satisfaction, its effects are typically overstated.&lt;&#x2F;p&gt;
&lt;p&gt;On a more general note, the truth is that we humans are not very good at predicting what makes us happy and fulfilled. In future posts, I will use microdata from large-scale household surveys to explore the determinants of individual wellbeing, and discuss how simple frameworks can be used to inform and design better policies (and, perhaps more importantly, to avoid bad ones).&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;!-- &lt;blockquote class=&quot;callout tip&quot;&gt;
    
    &lt;div class=&quot;icon&quot;&gt;
        &lt;svg xmlns=&quot;http:&#x2F;&#x2F;www.w3.org&#x2F;2000&#x2F;svg&quot; viewBox=&quot;0 0 24 24&quot; width=&quot;20&quot; height=&quot;20&quot;&gt;&lt;path d=&quot;M9.97308 18H11V13H13V18H14.0269C14.1589 16.7984 14.7721 15.8065 15.7676 14.7226C15.8797 14.6006 16.5988 13.8564 16.6841 13.7501C17.5318 12.6931 18 11.385 18 10C18 6.68629 15.3137 4 12 4C8.68629 4 6 6.68629 6 10C6 11.3843 6.46774 12.6917 7.31462 13.7484C7.40004 13.855 8.12081 14.6012 8.23154 14.7218C9.22766 15.8064 9.84103 16.7984 9.97308 18ZM10 20V21H14V20H10ZM5.75395 14.9992C4.65645 13.6297 4 11.8915 4 10C4 5.58172 7.58172 2 12 2C16.4183 2 20 5.58172 20 10C20 11.8925 19.3428 13.6315 18.2443 15.0014C17.624 15.7748 16 17 16 18.5V21C16 22.1046 15.1046 23 14 23H10C8.89543 23 8 22.1046 8 21V18.5C8 17 6.37458 15.7736 5.75395 14.9992Z&quot; fill=&quot;currentColor&quot;&gt;&lt;&#x2F;path&gt;&lt;&#x2F;svg&gt;
    &lt;&#x2F;div&gt;
    &lt;div class=&quot;content&quot;&gt;
        
        &lt;p&gt;&lt;strong&gt;Code and data&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
        
        &lt;p&gt;All the code used to produce the analyses in my blog posts (if any) is publicly available in my &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;pablogguz&#x2F;blog_posts_code&quot;&gt;blog&#x27;s code kitchen&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;

    &lt;&#x2F;div&gt;
&lt;&#x2F;blockquote&gt; --&gt;
&lt;hr &#x2F;&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;1&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;1&lt;&#x2F;sup&gt;
&lt;p&gt;A Likert-type scale is a psychometric scale commonly involved in research that employs questionnaires. It is the most widely used approach to scaling responses in survey research, such that the term is often used interchangeably with rating scale, although there are other types of rating scales. When responding to a Likert item, respondents specify their level of agreement or disagreement on a symmetric agree-disagree scale for a series of statements. Thus, the range captures the intensity of their feelings for a given item. For example, a 5-point Likert item could have the following options: 1. Strongly disagree, 2. Disagree, 3. Neither agree nor disagree, 4. Agree, 5. Strongly agree.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;2&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;2&lt;&#x2F;sup&gt;
&lt;p&gt;For instance, see &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;doi.org&#x2F;10.1016&#x2F;j.jebo.2010.09.016&quot;&gt;Di Tella et al. (2010)&lt;&#x2F;a&gt; and &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.aeaweb.org&#x2F;articles?id=10.1257&#x2F;aer.101.5.2226&quot;&gt;Kuhn et al. (2011)&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;br&gt;
&lt;h1 id=&quot;references&quot;&gt;References&lt;a class=&quot;zola-anchor&quot; href=&quot;#references&quot; aria-label=&quot;Anchor link for: references&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;ul&gt;
&lt;li&gt;Blanchflower, D. G., &amp;amp; Bryson, A. (2023). Were Covid and the Great Recession wellbeing Reducing? IZA Discussion Paper No. 16355.&lt;&#x2F;li&gt;
&lt;li&gt;Card, D., Mas, A., Moretti, E., &amp;amp; Saez, E. (2012). Inequality at Work: The Effect of Peer Salaries on Job Satisfaction. The American Economic Review, 102(6), 2981–3003. http:&#x2F;&#x2F;www.jstor.org&#x2F;stable&#x2F;41724678&lt;&#x2F;li&gt;
&lt;li&gt;Di Tella, R., Haisken-De New, J., &amp;amp; MacCulloch, R. (2010). Happiness adaptation to income and to status in an individual panel. Journal of Economic Behavior &amp;amp; Organization, 76(3), 834-852. https:&#x2F;&#x2F;doi.org&#x2F;10.1016&#x2F;j.jebo.2010.09.016&lt;&#x2F;li&gt;
&lt;li&gt;Easterlin, R. A., &amp;amp; O’Connor, K. J. (2022). Explaining happiness trends in Europe. Proceedings of the National Academy of Sciences, 119(37), e2210639119.&lt;&#x2F;li&gt;
&lt;li&gt;Kahneman, D., &amp;amp; Deaton, A. (2010). High income improves evaluation of life but not emotional wellbeing. Proceedings of the national academy of sciences, 107(38), 16489-16493.&lt;&#x2F;li&gt;
&lt;li&gt;Killingsworth, M.A. (2021). Experienced wellbeing rises with income, even above $75,000 per year. Proc. Natl. Acad. Sci. U.S.A. 118, e2016976118. https:&#x2F;&#x2F;doi.org&#x2F;10.1073&#x2F;pnas.2016976118&lt;&#x2F;li&gt;
&lt;li&gt;Killingsworth, M.A., Kahneman, D., &amp;amp; Sellers, B. (2023). Income and emotional wellbeing: A conflict resolved. Proc. Natl. Acad. Sci. https:&#x2F;&#x2F;doi.org&#x2F;10.1073&#x2F;pnas.2208661120&lt;&#x2F;li&gt;
&lt;li&gt;Kuhn, P., Kooreman, P., Soetevent, A., &amp;amp; Kapteyn, A. (2011). The Effects of Lottery Prizes on Winners and Their Neighbors: Evidence from the Dutch Postcode Lottery. American Economic Review, 101(5), 2226–2247. doi:10.1257&#x2F;aer.101.5.2226 &lt;&#x2F;li&gt;
&lt;li&gt;Lindqvist, E., Östling, R., &amp;amp; Cesarini, D. (2020). Long-run effects of lottery wealth on psychological wellbeing. The Review of Economic Studies, 87(6), 2703-2726. https:&#x2F;&#x2F;doi.org&#x2F;10.1093&#x2F;restud&#x2F;rdaa006&lt;&#x2F;li&gt;
&lt;li&gt;Perez-Truglia, Ricardo. 2020. &quot;The Effects of Income Transparency on wellbeing: Evidence from a Natural Experiment.&quot; American Economic Review, 110 (4): 1019-54.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
</content>
	</entry>
	<entry xml:lang="en">
		<title>Predicting socio-economic outcomes with Google Trends</title>
		<published>2024-03-10T00:00:00+00:00</published>
		<updated>2024-03-10T00:00:00+00:00</updated>
		<link href="https://pablogguz.github.io/blog/gtrends-badbunny/" type="text/html"/>
		<id>https://pablogguz.github.io/blog/gtrends-badbunny/</id>
		<content type="html">&lt;img src=&quot;&#x2F;img&#x2F;dalle_gtrends.webp&quot; width=&quot;500&quot;&#x2F;&gt;
&lt;br&gt;
&lt;p&gt;Everybody lies (&lt;em&gt;sometimes&lt;&#x2F;em&gt;).&lt;&#x2F;p&gt;
&lt;p&gt;This is both a universal truth and the title of a book by &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;http:&#x2F;&#x2F;sethsd.com&#x2F;&quot;&gt;Seth Stephens-Davidowitz&lt;&#x2F;a&gt;, former data scientist at Google and PhD in economics from Harvard. &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.amazon.co.uk&#x2F;Everybody-Lies-Internet-about-Really&#x2F;dp&#x2F;0062390856&quot;&gt;&lt;em&gt;Everybody lies&lt;&#x2F;em&gt;&lt;&#x2F;a&gt; is an incredibly fun read offering a compelling argument: that our online searches can reveal more about our true selves than we might like to admit. The book is not only a testament to the power of non-traditional data sources, but also an invitation to explore human behaviour through a different lens.&lt;&#x2F;p&gt;
&lt;p&gt;Stephens-Davidowitz&#x27;s work shows that Google searches can be used to predict a wide range of socio-economic outcomes when conventional data sources are not available or are contaminated by social desirability bias. The latter occurs when people are reluctant to disclose their true preferences or beliefs in surveys, which is particularly important when studying sensitive topics such as racial animus, sexual orientation, or mental health. The idea is quite simple: if revealing what you really think leads to social stigma, then you would be more likely to confess to Google than to a survey interviewer. In a sense, the unfiltered truth of humanity lies in Google&#x27;s autocomplete bar.&lt;&#x2F;p&gt;
&lt;p&gt;In &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;people.cs.umass.edu&#x2F;~brenocon&#x2F;smacss2015&#x2F;papers&#x2F;StephensDawidowitz2014.pdf&quot;&gt;one of his academic papers&lt;&#x2F;a&gt; published in the &lt;em&gt;Journal of Public Economics&lt;&#x2F;em&gt;, Stephens-Davidowitz uses Google searches to measure racial animus in the U.S. and its impact on the 2008 and 2012 presidential elections – specifically, he finds that racially charged Google search rates are a strong negative predictor of Obama&#x27;s vote share. The point estimates in the preferred specification suggest that racial animus cost Obama around 4 percentage points in the 2008 election, which is significantly above previous estimates based on survey data.&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;div id=&quot;table1&quot;&gt;
&lt;figcaption&gt;Table 1: Google search terms and socio-economic outcomes &lt;&#x2F;figcaption&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Term&lt;&#x2F;th&gt;&lt;th&gt;Underlying variable&lt;&#x2F;th&gt;&lt;th&gt;t-Stat&lt;&#x2F;th&gt;&lt;th&gt;$R^2$&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;God&lt;&#x2F;td&gt;&lt;td&gt;Percent believe in god&lt;&#x2F;td&gt;&lt;td&gt;8.45&lt;&#x2F;td&gt;&lt;td&gt;0.65&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Gun&lt;&#x2F;td&gt;&lt;td&gt;Percent own gun&lt;&#x2F;td&gt;&lt;td&gt;8.94&lt;&#x2F;td&gt;&lt;td&gt;0.62&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;African American(s)&lt;&#x2F;td&gt;&lt;td&gt;Percent Black&lt;&#x2F;td&gt;&lt;td&gt;13.15&lt;&#x2F;td&gt;&lt;td&gt;0.78&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Hispanic&lt;&#x2F;td&gt;&lt;td&gt;Percent Hispanic&lt;&#x2F;td&gt;&lt;td&gt;8.71&lt;&#x2F;td&gt;&lt;td&gt;0.61&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Jewish&lt;&#x2F;td&gt;&lt;td&gt;Percent Jewish&lt;&#x2F;td&gt;&lt;td&gt;17.08&lt;&#x2F;td&gt;&lt;td&gt;0.86&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;p class=&quot;figure-note&quot;&gt;Source: &lt;a href=&quot;https:&#x2F;&#x2F;people.cs.umass.edu&#x2F;~brenocon&#x2F;smacss2015&#x2F;papers&#x2F;StephensDawidowitz2014.pdf&quot;&gt;Stephens-Davidowitz (2015)&lt;&#x2F;a&gt;, Table 1. The t-stat and $R^2$ are from a regression with the normalized search volume of the
word(s) in the first column as the independent variable and measures of the value in
the second column as the dependent variable. The normalized search volume for all
terms is from 2004 to 2007. All data are at the state level.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;br&gt;
&lt;p&gt;But the potential of Google searches goes beyond measuring racial animus. In &lt;a href=&quot;#table1&quot;&gt;Table 1&lt;&#x2F;a&gt;, I show some of Stephen-Davidowitz&#x27;s examples on how Google searches can be exploited to predict various population-level outcomes. For example, Google searches for the word &quot;Gun&quot; explain 62% of the variation in states&#x27; gun ownership rates, and Google searches for the word &quot;Hispanic&quot; explain 61% of the variation in the percentage of Hispanic population across U.S. states.&lt;&#x2F;p&gt;
&lt;p&gt;Can we do better? In this post, I will show you how we can leverage the popularity of Latin urban music artists in Google searches to predict state-level Hispanic population shares in the U.S.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;some-background-on-google-trends-data&quot;&gt;Some background on Google Trends data&lt;a class=&quot;zola-anchor&quot; href=&quot;#some-background-on-google-trends-data&quot; aria-label=&quot;Anchor link for: some-background-on-google-trends-data&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;Data on Google searches is available through &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;trends.google.com&#x2F;trends&#x2F;&quot;&gt;Google Trends&lt;&#x2F;a&gt;, a tool that allows users to explore the popularity of search terms over time and across different regions within countries. The data is normalized to the time and location of a query, so that the popularity of a search term is measured relative to the total number of searches in a given region and time period. In other words, Google Trends gives measures of relative search term popularity, not an absolute measure of the number of searches for a given term.[^1] Relative popularity is measured on a scale from 0 to 100, where 100 is the most popular point in time or location for a given search term. For example, value of 50 means that the term is at half its peak observed popularity.&lt;&#x2F;p&gt;
&lt;p&gt;There are a bazillion online resources on how to use Google Trends, so I won&#x27;t go into the details here. As a starting point, I would recommend you to check out the &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;newsinitiative.withgoogle.com&#x2F;es-es&#x2F;resources&#x2F;trainings&#x2F;basics-of-google-trends&#x2F;&quot;&gt;Google Trends basics page&lt;&#x2F;a&gt; and the &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;support.google.com&#x2F;trends&#x2F;answer&#x2F;4359582?hl=en&quot;&gt;search tips page&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;predicting-hispanic-population-shares-with-google-trends&quot;&gt;Predicting Hispanic population shares with Google Trends&lt;a class=&quot;zola-anchor&quot; href=&quot;#predicting-hispanic-population-shares-with-google-trends&quot; aria-label=&quot;Anchor link for: predicting-hispanic-population-shares-with-google-trends&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;We will combine Google Trends data on the popularity of search terms related to Latin urban music superstars with data on the Hispanic population shares across U.S. states sourced from the &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.census.gov&#x2F;programs-surveys&#x2F;acs&#x2F;data.html&quot;&gt;American Community Survey (ACS)&lt;&#x2F;a&gt; microdata for the year 2022.[^2]&lt;&#x2F;p&gt;
&lt;p&gt;First, we will use the search term &quot;bad bunny&quot; to predict Hispanic population shares. &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;open.spotify.com&#x2F;intl-es&#x2F;artist&#x2F;4q3ewBCX7sLwd24euuV69X&quot;&gt;Bad Bunny&lt;&#x2F;a&gt; is arguably the most popular Latin artist worldwide nowadays – in 2022, he was Billboard&#x27;s top artist of the year and the &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.billboard.com&#x2F;pro&#x2F;bad-bunny-top-streaming-artist-year-end-charts-2022&#x2F;&quot;&gt;most streamed artist&lt;&#x2F;a&gt; in Spotify and Apple Music. He also has a &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.instagram.com&#x2F;badbunnypr&#x2F;&quot;&gt;huge following&lt;&#x2F;a&gt; on Instagram with over 40 million followers, and its popularity has played a crucial role in the reivindication of the Latino culture and identity in the U.S. (&lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=QdQEljUMCEM&quot;&gt;&lt;em&gt;ahora todos quieren ser latinos&lt;&#x2F;em&gt;&lt;&#x2F;a&gt;).&lt;&#x2F;p&gt;
&lt;p&gt;Let&#x27;s do it. To start with, I show the Hispanic population shares across U.S. states in 2022 in &lt;a href=&quot;#fig1&quot;&gt;Figure 1&lt;&#x2F;a&gt;. Unsurprisingly, the Hispanic population is concentrated in the South and West regions of the U.S. The states with the highest Hispanic population shares are New Mexico, California, and Texas, with shares of 50%, 40.3%, and 40.2%, respectively. The states with the lowest Hispanic population shares are West Virginia, Maine, and Vermont, with shares of around 2% each.&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;figure id=&quot;fig1&quot;&gt;
  &lt;figcaption&gt;Figure 1&lt;&#x2F;figcaption&gt;
  &lt;img src=&quot;&#x2F;img&#x2F;map_hispanic.png&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; border: 0px none;&quot;&gt;&lt;&#x2F;iframe&gt;
&lt;&#x2F;figure&gt;
&lt;!-- &lt;p class=&quot;figure-note&quot;&gt;Source: 2022 ACS.&lt;&#x2F;a&gt;&lt;&#x2F;p&gt; --&gt;
&lt;br&gt;
&lt;p&gt;Now, let&#x27;s do the same for the popularity of the search term &quot;bad bunny&quot; in the U.S. in 2022:&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;figure id=&quot;fig2&quot;&gt;
  &lt;figcaption&gt;Figure 2&lt;&#x2F;figcaption&gt;
  &lt;img src=&quot;&#x2F;img&#x2F;map_badbunny.png&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; border: 0px none;&quot;&gt;&lt;&#x2F;iframe&gt;
&lt;&#x2F;figure&gt;
&lt;!-- &lt;p class=&quot;figure-note&quot;&gt;Source: 2022 ACS.&lt;&#x2F;a&gt;&lt;&#x2F;p&gt; --&gt;
&lt;br&gt;
&lt;p&gt;Similar, right?&lt;&#x2F;p&gt;
&lt;p&gt;We can visualize the relationship between the popularity of the search term &quot;bad bunny&quot; and the Hispanic population shares across U.S. states in a scatter plot (see &lt;a href=&quot;#fig3&quot;&gt;Figure 3&lt;&#x2F;a&gt;). By running a simple OLS regression of the Hispanic population share on the popularity of the search term, we get that 76% of the variation in the Hispanic population share across U.S. states can be explained by the relative popularity of Bad Bunny in Google searches. This implies a 15 percentage point increase in the share of variation explained relative to just using the term &quot;Hispanic&quot; as a predictor.&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;figure id=&quot;fig3&quot;&gt;
  &lt;figcaption&gt;Figure 3&lt;&#x2F;figcaption&gt;
  &lt;img src=&quot;&#x2F;img&#x2F;corr_badbunny_hispanic.png&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; border: 0px none;&quot;&gt;&lt;&#x2F;iframe&gt;
&lt;&#x2F;figure&gt;
&lt;!-- &lt;p class=&quot;figure-note&quot;&gt;Source: 2022 ACS.&lt;&#x2F;a&gt;&lt;&#x2F;p&gt; --&gt;
&lt;br&gt;
&lt;p&gt;There is one state that is particularly off path: New Mexico (if we were to omit it, then the share of variation explained would go up to 89%). New Mexico has a Hispanic population share of 50%, but the popularity of Bad Bunny in Google searches is only 56% of the most popular point location for the term.&lt;&#x2F;p&gt;
&lt;p&gt;One particularity of New Mexico is that it has the highest population share with Mexican origin in the U.S. What if we focused on the popularity of a Mexican artist, and looked at Mexican population shares instead? I got curious about this, so I replicated the analysis exploiting the popularity of the search term &quot;peso pluma&quot;. &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;open.spotify.com&#x2F;intl-es&#x2F;artist&#x2F;12GqGscKJx3aE4t07u7eVZ&quot;&gt;Peso Pluma&lt;&#x2F;a&gt; is a Mexican artist whose popularity has skyrocketed during the last couple of years, being &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.billboard.com&#x2F;espanol&#x2F;noticias&#x2F;peso-pluma-artista-mas-visto-youtube-2023-estados-unidos-1235563630&#x2F;&quot;&gt;Youtube&#x27;s most viewed artist of 2023 in the U.S.&lt;&#x2F;a&gt;&lt;&#x2F;p&gt;
&lt;p&gt;The results are shown in &lt;a href=&quot;#fig4&quot;&gt;Figure 4&lt;&#x2F;a&gt;. The fit is breathtaking: Peso Pluma&#x27;s popularity in Google searches explains 94% of the variation in the Mexican population share across U.S. states.&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;figure id=&quot;fig4&quot;&gt;
  &lt;figcaption&gt;Figure 4&lt;&#x2F;figcaption&gt;
  &lt;img src=&quot;&#x2F;img&#x2F;corr_pesopluma_mexican.png&quot; loading=&quot;lazy&quot; style=&quot;width: 100%; border: 0px none;&quot;&gt;&lt;&#x2F;iframe&gt;
&lt;&#x2F;figure&gt;
&lt;!-- &lt;p class=&quot;figure-note&quot;&gt;Source: 2022 ACS.&lt;&#x2F;a&gt;&lt;&#x2F;p&gt; --&gt;
&lt;br&gt;
&lt;h1 id=&quot;measuring-what-matters&quot;&gt;Measuring what matters&lt;a class=&quot;zola-anchor&quot; href=&quot;#measuring-what-matters&quot; aria-label=&quot;Anchor link for: measuring-what-matters&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;Fine, this is cute. But what&#x27;s the point?&lt;&#x2F;p&gt;
&lt;p&gt;The point is that Google searches can be used to measure things that actually matter. We might not need to predict historical Hispanic population shares, or for that matter any type of socio-economic outcome that can be easily derived from traditional data sources. But we might be interested, for instance, in nowcasting suicide rates or the prevalence of mental health disorders. Insofar as these figures could guide the allocation of resources in healthcare systems and provide early warning signs for potential crises, they could be a quite important estimand for policy.&lt;&#x2F;p&gt;
&lt;p&gt;In &lt;a href=&quot;#table1&quot;&gt;Table 2&lt;&#x2F;a&gt;, I illustrate the comparison between how state-level suicide rates in the U.S. correlate with both Google search terms and self-reported measures of suicide risk. While Google searches are not perfect, they prove to be more reliable than the traditional survey-derived data, which yields statistically insignificant correlations.&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;div id=&quot;table1&quot;&gt;
&lt;figcaption&gt;Table 2: Correlations between Google search terms and suicide rates in the U.S. &lt;&#x2F;figcaption&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th style=&quot;text-align: left&quot;&gt;Predictor&lt;&#x2F;th&gt;&lt;th&gt;$r$&lt;&#x2F;th&gt;&lt;th&gt;$p$&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;Google search term&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;&lt;&#x2F;td&gt;&lt;td&gt;&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: left&quot;&gt;Suicide&lt;&#x2F;td&gt;&lt;td&gt;-.06&lt;&#x2F;td&gt;&lt;td&gt;.63&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: left&quot;&gt;How to suicide&lt;&#x2F;td&gt;&lt;td&gt;.49&lt;&#x2F;td&gt;&lt;td&gt;&amp;lt; .001&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: left&quot;&gt;How to kill yourself&lt;&#x2F;td&gt;&lt;td&gt;.61&lt;&#x2F;td&gt;&lt;td&gt;&amp;lt; .001&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: left&quot;&gt;How to commit suicide&lt;&#x2F;td&gt;&lt;td&gt;.63&lt;&#x2F;td&gt;&lt;td&gt;&amp;lt; .001&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: left&quot;&gt;Painless suicide&lt;&#x2F;td&gt;&lt;td&gt;.59&lt;&#x2F;td&gt;&lt;td&gt;&amp;lt; .001&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: left&quot;&gt;Composite of all four suicidal terms&lt;&#x2F;td&gt;&lt;td&gt;.67&lt;&#x2F;td&gt;&lt;td&gt;&amp;lt; .001&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;Self-report measure&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;&lt;&#x2F;td&gt;&lt;td&gt;&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: left&quot;&gt;Suicidal thoughts&lt;&#x2F;td&gt;&lt;td&gt;-.17&lt;&#x2F;td&gt;&lt;td&gt;.22&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: left&quot;&gt;Suicidal plans&lt;&#x2F;td&gt;&lt;td&gt;-.14&lt;&#x2F;td&gt;&lt;td&gt;.30&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: left&quot;&gt;Suicidal attempts&lt;&#x2F;td&gt;&lt;td&gt;-.03&lt;&#x2F;td&gt;&lt;td&gt;.82&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: left&quot;&gt;Composite of self-reported suicidality&lt;&#x2F;td&gt;&lt;td&gt;-.16&lt;&#x2F;td&gt;&lt;td&gt;.25&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;p class=&quot;figure-note&quot;&gt;Source: &lt;a href=&quot;https:&#x2F;&#x2F;journals.sagepub.com&#x2F;doi&#x2F;abs&#x2F;10.1177&#x2F;2167702615593475&quot;&gt;Ma-Kellams et al. (2016)&lt;&#x2F;a&gt;, Table 1. Data for 2008 and 2009. The column $r$ shows the correlation between suicidal death rates and the relevant Google search term at the state level. Self-reported measures are sourced from the National Survey on Drug Use and Health.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;br&gt;
&lt;p&gt;We can also use Google searches to track economic activity in real time. For example, the OECD has developed a &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.oecd.org&#x2F;economy&#x2F;weekly-tracker-of-gdp-growth&#x2F;&quot;&gt;tool&lt;&#x2F;a&gt; that uses Google searches to nowcast GDP growth on a weekly basis. The algorithm gathers data related to various realms of economic activity, using search terms related to consumption (like &quot;vehicles&quot; and &quot;household appliances&quot;), labour market conditions (such as &quot;unemployment benefits&quot;), the housing sector (like &quot;real estate agencies&quot; and &quot;mortgages&quot;), business-related services, industrial activity, trade, and economic sentiment, among others. So, next time you are casually checking (again) whether rent prices are going up in your area, you might be contributing to a fancy economic forecast.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;wrapping-up&quot;&gt;Wrapping up&lt;a class=&quot;zola-anchor&quot; href=&quot;#wrapping-up&quot; aria-label=&quot;Anchor link for: wrapping-up&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;In this post, I have shown some examples on how Google searches can be used to proxy for a wide range of outcomes when conventional data sources are not available or are contaminated by social desirability bias. This often involves being creative and leveraging pop-culture knowledge – as we saw, Bad Bunny helps us predict Hispanic population shares better than the term &quot;Hispanic&quot; itself.&lt;&#x2F;p&gt;
&lt;p&gt;That said, it does not always work. Using Google Trends for research purposes can be tricky, as most of the time you will lack the granularity needed to link the data to the relevant population. In the U.S., the data is only available at the state and metro area levels, and it is not possible to obtain data for finer geographic areas or specific demographic groups.[^3] For some countries, the minimum level of aggregation is too coarse to be useful: for instance, the UK&#x27;s finest administrative boundary in Google Trends corresponds to the &lt;em&gt;region&lt;&#x2F;em&gt; level, which means that you will not have variation within England, Scotland, Wales, or Northern Ireland.&lt;&#x2F;p&gt;
&lt;p&gt;Still, I think it&#x27;s quite cool. I will leave below a bunch of potential applications for Google Trends data that I would find particularly interesting:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Sexual minorities&lt;&#x2F;strong&gt;: can we use Google searches to estimate sexual minority population shares in settings where sexual orientation and identity are heavily stigmatized? Stephens-Davidowitz has some work on this for the U.S., but I would be curious to see how this could be extended to countries where sexual minorities are criminalized.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Political outcomes&lt;&#x2F;strong&gt;: there has been a rising concern about how polls often underestimate voting intention for far-right parties (for instance, see &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.aljazeera.com&#x2F;news&#x2F;2022&#x2F;10&#x2F;6&#x2F;how-did-brazils-pollsters-underestimate-support-for-bolsonaro&quot;&gt;here&lt;&#x2F;a&gt;). Could we predict vote shares for controversial candidates at scale with Google searches?&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Emergency situations and public safety&lt;&#x2F;strong&gt;: during emergencies, real-time data from Google Trends can provide insights into public sentiment or areas of concern, which might be faster or more accurate other methods. For instance, in the context of natural disasters, spikes in searches for emergency services or relief efforts can indicate the most affected areas or the public’s most urgent needs.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;I will leave it here. If you have worked with Google Trends data before and have some cool applications in mind you would like to share, please feel free to get in touch!&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;!-- &lt;blockquote class=&quot;callout tip&quot;&gt;
    
    &lt;div class=&quot;icon&quot;&gt;
        &lt;svg xmlns=&quot;http:&#x2F;&#x2F;www.w3.org&#x2F;2000&#x2F;svg&quot; viewBox=&quot;0 0 24 24&quot; width=&quot;20&quot; height=&quot;20&quot;&gt;&lt;path d=&quot;M9.97308 18H11V13H13V18H14.0269C14.1589 16.7984 14.7721 15.8065 15.7676 14.7226C15.8797 14.6006 16.5988 13.8564 16.6841 13.7501C17.5318 12.6931 18 11.385 18 10C18 6.68629 15.3137 4 12 4C8.68629 4 6 6.68629 6 10C6 11.3843 6.46774 12.6917 7.31462 13.7484C7.40004 13.855 8.12081 14.6012 8.23154 14.7218C9.22766 15.8064 9.84103 16.7984 9.97308 18ZM10 20V21H14V20H10ZM5.75395 14.9992C4.65645 13.6297 4 11.8915 4 10C4 5.58172 7.58172 2 12 2C16.4183 2 20 5.58172 20 10C20 11.8925 19.3428 13.6315 18.2443 15.0014C17.624 15.7748 16 17 16 18.5V21C16 22.1046 15.1046 23 14 23H10C8.89543 23 8 22.1046 8 21V18.5C8 17 6.37458 15.7736 5.75395 14.9992Z&quot; fill=&quot;currentColor&quot;&gt;&lt;&#x2F;path&gt;&lt;&#x2F;svg&gt;
    &lt;&#x2F;div&gt;
    &lt;div class=&quot;content&quot;&gt;
        
        &lt;p&gt;&lt;strong&gt;Code and data&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
        
        &lt;p&gt;All the code used to produce the analyses in my blog posts (if any) is publicly available in my &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;pablogguz&#x2F;blog_posts_code&quot;&gt;blog&#x27;s code kitchen&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;

    &lt;&#x2F;div&gt;
&lt;&#x2F;blockquote&gt; --&gt;
&lt;hr &#x2F;&gt;
&lt;p&gt;[^1] However, with some assumptions and creativity, it is possible to come up with proxies for absolute Google search volumnes. See Appendix C in &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;people.cs.umass.edu&#x2F;~brenocon&#x2F;smacss2015&#x2F;papers&#x2F;StephensDawidowitz2014.pdf&quot;&gt;Stephens-Davidowitz (2015)&lt;&#x2F;a&gt; for more details.&lt;&#x2F;p&gt;
&lt;p&gt;[^2] I identify people with Hispanic&#x2F;Latino origin in the ACS using the variable &lt;code&gt;HISPAN&lt;&#x2F;code&gt;. More details can be found &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;usa.ipums.org&#x2F;usa-action&#x2F;variables&#x2F;HISPAN#description_section&quot;&gt;here&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;[^3] One further caveat is that the Google Trends metro areas are custom aggregates and do not correspond to standard metropolitan statistical areas (MSAs) or metropolitan divisions. Due to this, there has been some attempts to map Google Trends metro areas to more granular geographic units, such as &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;sites.google.com&#x2F;view&#x2F;jacob-schneider&#x2F;resources&quot;&gt;counties in the U.S&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;br&gt;
&lt;h1 id=&quot;references&quot;&gt;References&lt;a class=&quot;zola-anchor&quot; href=&quot;#references&quot; aria-label=&quot;Anchor link for: references&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;ul&gt;
&lt;li&gt;Stephens-Davidowitz, S. (2017). &lt;em&gt;Everybody lies: Big data, new data, and what the internet can tell us about who we really are&lt;&#x2F;em&gt;. Bloomsbury Publishing.&lt;&#x2F;li&gt;
&lt;li&gt;Stephens-Davidowitz, S. (2015). The cost of racial animus on a black candidate: Evidence using Google search data. &lt;em&gt;Journal of Public Economics&lt;&#x2F;em&gt;, 118, 26-40.&lt;&#x2F;li&gt;
&lt;li&gt;Ma-Kellams, C., Or, F., Baek, J. H., &amp;amp; Kawachi, I. (2016). Rethinking suicide surveillance: Google search data and self-reported suicidality differentially estimate completed suicide risk. Clinical Psychological Science, 4(3), 480-484.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
</content>
	</entry>
	<entry xml:lang="en">
		<title>A modern IDE for Stata with VSCode + Quarto</title>
		<published>2024-02-27T00:00:00+00:00</published>
		<updated>2024-02-27T00:00:00+00:00</updated>
		<link href="https://pablogguz.github.io/blog/stata-vscode-quarto/" type="text/html"/>
		<id>https://pablogguz.github.io/blog/stata-vscode-quarto/</id>
		<content type="html">&lt;blockquote class=&quot;callout alert&quot;&gt;
    
    &lt;div class=&quot;icon&quot;&gt;
        &lt;svg xmlns=&quot;http:&#x2F;&#x2F;www.w3.org&#x2F;2000&#x2F;svg&quot; viewBox=&quot;0 0 24 24&quot; width=&quot;20&quot; height=&quot;20&quot;&gt;&lt;path d=&quot;M4.00098 20V14C4.00098 9.58172 7.5827 6 12.001 6C16.4193 6 20.001 9.58172 20.001 14V20H21.001V22H3.00098V20H4.00098ZM6.00098 20H18.001V14C18.001 10.6863 15.3147 8 12.001 8C8.68727 8 6.00098 10.6863 6.00098 14V20ZM11.001 2H13.001V5H11.001V2ZM19.7792 4.80761L21.1934 6.22183L19.0721 8.34315L17.6578 6.92893L19.7792 4.80761ZM2.80859 6.22183L4.22281 4.80761L6.34413 6.92893L4.92991 8.34315L2.80859 6.22183ZM7.00098 14C7.00098 11.2386 9.23956 9 12.001 9V11C10.3441 11 9.00098 12.3431 9.00098 14H7.00098Z&quot; fill=&quot;currentColor&quot;&gt;&lt;&#x2F;path&gt;&lt;&#x2F;svg&gt;
    &lt;&#x2F;div&gt;
    &lt;div class=&quot;content&quot;&gt;
        
        &lt;p&gt;&lt;strong&gt;Update Aug 2024&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
        
        &lt;p&gt;I only used this setup for a couple of weeks before switching to a non-interactive integration that allows for running &lt;code&gt;.do&lt;&#x2F;code&gt; files within VSCode. The lack of a data browser and the variable explorer in the Jupyter kernel was a dealbreaker for me (plus I have grown to find interactive notebooks a bit annoying except for very specific tasks). To set this up, follow the instructions &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;wulizyk&#x2F;Notebook_Stata&#x2F;wiki&#x2F;Run-STATA-in-VScode&quot;&gt;here&lt;&#x2F;a&gt; under the &quot;way one&quot; section.&lt;&#x2F;p&gt;
&lt;p&gt;Moreover, I have increasingly found myself using R for most of my work over the last year, and I have to stay that I no longer agree with some of the points I made in this post. Most of the time, the so-claimed &quot;superiority&quot; of Stata for reduced-form empirical work is just a skill issue. My advice would be to invest time in mastering R or Python, and &lt;em&gt;then&lt;&#x2F;em&gt; decide whether you really need to use Stata (you probably won&#x27;t).&lt;&#x2F;p&gt;

    &lt;&#x2F;div&gt;
&lt;&#x2F;blockquote&gt;&lt;img src=&quot;&#x2F;img&#x2F;dalle_vscode_stata_v2.webp&quot; width=&quot;600&quot;&#x2F;&gt;
&lt;p&gt;Over the last few decades, applied microeconomic research has witnessed what is widely termed as &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.aeaweb.org&#x2F;articles?id=10.1257&#x2F;jep.24.2.3&quot;&gt;the credibility revolution&lt;&#x2F;a&gt;, a paradigm shift that emphasized the importance of experimental and quasi-experimental methods to credibly identify causal relationships between variables. This transformative period has not only influenced the methodologies employed in academic papers but also strengthened the field&#x27;s impact on policy-making, by providing more solid and rigorous evidence upon which to base decisions.&lt;&#x2F;p&gt;
&lt;p&gt;This emphasis on methodologically robust empirical approaches has resulted in bigger, more complex projects, naturally elevating the importance of data skills within the field. However, despite the increasing importance of coding in economics, many economists lag in adopting modern coding practices and development environments. One of the most widely used tools in the profession, &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.stata.com&#x2F;&quot;&gt;Stata&lt;&#x2F;a&gt;, is not even a programming language but a purpose-built statistical package.&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#1&quot;&gt;1&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt; Stata is also commercial software, meaning that it does not have the benefits of open-source languages like Python or R, which have amassed extensive communities contributing to their development and offering support. Due to its nature, it has traditionally been hard to integrate Stata with modern IDEs, resulting in a less efficient workflow when one has to write code in multiple languages or use the latest tools for version control, code completion, and so on.&lt;&#x2F;p&gt;
&lt;p&gt;This post provides a quick tutorial on setting up a modern development environment for Stata with Visual Studio Code and Quarto. It is by no means perfect and it comes with its own caveats, but it is a step in the right direction for those who want to modernize a bit their workflow without having to give up on Stata.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;why-stata&quot;&gt;Why Stata&lt;a class=&quot;zola-anchor&quot; href=&quot;#why-stata&quot; aria-label=&quot;Anchor link for: why-stata&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;I know what you are thinking.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;jUsT dOn&#x27;t usE sTAta BrO&lt;&#x2F;strong&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;I get it. I code in multiple languages and it is indeed not a good idea to &lt;em&gt;only&lt;&#x2F;em&gt; use Stata, simply because there are lots of cool things (web scraping, geospatial stuff, ML) that you would miss out on by doing so.&lt;&#x2F;p&gt;
&lt;p&gt;However, the truth is that if you do applied microeconomics for a living, &lt;strong&gt;it would be very stupid not to use Stata&lt;&#x2F;strong&gt; sometimes. Whoever tells you otherwise is i) not an applied microeconomist (good for them!) or ii) underestimating how important Stata is in the profession:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Stata can be found in &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.r-bloggers.com&#x2F;2023&#x2F;12&#x2F;usage-shares-of-programming-languages-in-economics-research&#x2F;&quot;&gt;more than 70% of replication packages&lt;&#x2F;a&gt; for published papers in economics. It is followed by Matlab (which we deliberately ignore since it is aimed at structural work), and R comes in third with less than 10%. I am not going to argue whether this is a good or a bad equilibrium but &lt;em&gt;it is an equilibrium&lt;&#x2F;em&gt;, and if you want collaborate with others and be able to replicate their work, you will have to use Stata&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;For reduced-form econometrics, Stata is unparalleled. Period. I have been writing code for +3.5 years and it still amazes me how ridiculously cumbersome R and Python can be relative to Stata when it comes to running regressions. For some specifications, it is simply not possible to use something else unless you want to code the estimator from scratch (good luck running modern difference-in-differences estimators for heterogeneous treatment effects in Python, for example)&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;Economists love regression tables. And again, the functionalities and flexibility of &lt;code&gt;esttab&lt;&#x2F;code&gt; to export fully formatted and reproducible LaTeX tables in Stata are light years ahead of the alternatives in other languages (e.g., &lt;code&gt;stargazer&lt;&#x2F;code&gt; in R)&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;For most people doing applied micro research or policy work, &lt;strong&gt;coding will hardly ever be their comparative advantage&lt;&#x2F;strong&gt;. Instead, your value lies in the quality of your ideas and how well you execute them – coding is only useful insofar as it helps in that purpose. Some folks (myself included) enjoy venturing off the beaten path, but most researchers would prioritize research productivity if the marginal cost of learning new data skills is sufficiently high. In this sense, Stata has a quite shallow and gentle learning curve, so usually people learn it as their first &quot;language&quot; and then just stick to it. I am not saying this is good (I think is bad!), but in the end people respond to incentives&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h1 id=&quot;why-vscode&quot;&gt;Why VSCode&lt;a class=&quot;zola-anchor&quot; href=&quot;#why-vscode&quot; aria-label=&quot;Anchor link for: why-vscode&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;VSCode is Microsoft&#x27;s open-source code editor. It supports a wide array of programming languages and comes with the usual IDE features like syntax highlighting, code completion, &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;features&#x2F;copilot&quot;&gt;Copilot&lt;&#x2F;a&gt;, and version control integration.&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#copilot&quot;&gt;2&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;&lt;&#x2F;p&gt;
&lt;p&gt;More importantly, turns out that it is fairly straightforward to integrate VSCode with Stata via plug-ins. VSCode also has nice LaTeX extensions, which in my opinion makes it the only real alternative to Overleaf. This means that you can seamlessly integrate all your workflow into a single code editor, from data cleaning to drafting.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;why-quarto&quot;&gt;Why Quarto&lt;a class=&quot;zola-anchor&quot; href=&quot;#why-quarto&quot; aria-label=&quot;Anchor link for: why-quarto&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;Quarto is an open-source scientific and technical publishing system that allows users to create dynamic documents, slide decks, and reports in various formats from plain text files. Quarto files support multiple programming languages and can be rendered into interactive HTML documents, which allows the user to combine code execution with high-quality document production.&lt;&#x2F;p&gt;
&lt;p&gt;The main advantage of using a notebook-based interface relative to the native Stata editor is that it allows you to create and present exploratory analysis combining code, output, and text (including comments, LaTeX equations, and more), all in one convenient place.&lt;&#x2F;p&gt;
&lt;p&gt;And last but not least, Quarto outputs are &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;quarto.org&#x2F;docs&#x2F;gallery&#x2F;#interactive-docs&quot;&gt;beautiful&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;setup&quot;&gt;Setup&lt;a class=&quot;zola-anchor&quot; href=&quot;#setup&quot; aria-label=&quot;Anchor link for: setup&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;h2 id=&quot;step-1-install-quarto&quot;&gt;Step 1: Install Quarto&lt;a class=&quot;zola-anchor&quot; href=&quot;#step-1-install-quarto&quot; aria-label=&quot;Anchor link for: step-1-install-quarto&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;Go to &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;quarto.org&#x2F;&quot;&gt;https:&#x2F;&#x2F;quarto.org&#x2F;&lt;&#x2F;a&gt;, click on &quot;Get Started&quot;, download the latest Quarto version available, and install it on your machine.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;img src=&quot;&#x2F;img&#x2F;quarto_example.JPG&quot; alt=&quot;Alt text&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;h2 id=&quot;step-2-install-vscode&quot;&gt;Step 2: Install VSCode&lt;a class=&quot;zola-anchor&quot; href=&quot;#step-2-install-vscode&quot; aria-label=&quot;Anchor link for: step-2-install-vscode&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;Go to &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;code.visualstudio.com&#x2F;download&quot;&gt;https:&#x2F;&#x2F;code.visualstudio.com&#x2F;download&lt;&#x2F;a&gt;, download the latest VSCode version available, and install it on your machine.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;img src=&quot;&#x2F;img&#x2F;vscode_example.JPG&quot; alt=&quot;Alt text&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;h2 id=&quot;step-3-install-vscode-extensions&quot;&gt;Step 3: Install VSCode extensions&lt;a class=&quot;zola-anchor&quot; href=&quot;#step-3-install-vscode-extensions&quot; aria-label=&quot;Anchor link for: step-3-install-vscode-extensions&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;You will have to install two VSCode extensions:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;To get nice Stata syntax highlighting and code autocompletion, you will have to install the &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;marketplace.visualstudio.com&#x2F;items?itemName=kylebarron.stata-enhanced&quot;&gt;Stata Enhanced&lt;&#x2F;a&gt; package. Open VSCode, go to the extensions tab (highlighted in red in the screenshot), search for the extension, and install it.&lt;&#x2F;li&gt;
&lt;li&gt;For integrated render and preview for Quarto documents, repeat the steps above and install the &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;marketplace.visualstudio.com&#x2F;items?itemName=quarto.quarto&quot;&gt;Quarto&lt;&#x2F;a&gt; extension.&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;&lt;img src=&quot;&#x2F;img&#x2F;stataenhanced_example.jpg&quot; alt=&quot;Alt text&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;h2 id=&quot;step-4-install-nbstata&quot;&gt;Step 4: Install &lt;code&gt;nbstata&lt;&#x2F;code&gt;&lt;a class=&quot;zola-anchor&quot; href=&quot;#step-4-install-nbstata&quot; aria-label=&quot;Anchor link for: step-4-install-nbstata&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;hugetim.github.io&#x2F;nbstata&#x2F;&quot;&gt;&lt;code&gt;nbstata&lt;&#x2F;code&gt;&lt;&#x2F;a&gt; is a Jupyter kernel for Stata, built on top of the functionalities of the &lt;code&gt;pystata&lt;&#x2F;code&gt; Python package introduced in Stata version 17. This means that you will need Stata 17 or higher to install it.&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#3&quot;&gt;3&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;&lt;&#x2F;p&gt;
&lt;p&gt;You can find the full user and installation guide &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;hugetim.github.io&#x2F;nbstata&#x2F;user_guide.html&quot;&gt;here&lt;&#x2F;a&gt;. For the curious, I encourage you to go through the author&#x27;s guide. For the lazy ones, the short version goes as follows:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;To install &lt;code&gt;nbstata&lt;&#x2F;code&gt;, you need an installation of Python 3.7 or higher. For first-time Python users, I would follow the suggestion from the author of &lt;code&gt;nbstata&lt;&#x2F;code&gt; and just install the &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;www.anaconda.com&#x2F;download&#x2F;&quot;&gt;Anaconda distribution&lt;&#x2F;a&gt;. If you want a lighter, minimalistic version, you can go for &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;docs.anaconda.com&#x2F;free&#x2F;miniconda&#x2F;&quot;&gt;Miniconda&lt;&#x2F;a&gt;.&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#4&quot;&gt;4&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;&lt;&#x2F;li&gt;
&lt;li&gt;Once you have installed Anaconda, open the Anaconda prompt and type&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;pre class=&quot;giallo z-code&quot;&gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;pip install nbstata&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;ol start=&quot;3&quot;&gt;
&lt;li&gt;Finally, install &lt;code&gt;nbstata&lt;&#x2F;code&gt; by running&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;pre class=&quot;giallo z-code&quot;&gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;python -m nbstata.install&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;And that is it! Your machine should be now ready to render Quarto Markdown (&lt;code&gt;.qmd&lt;&#x2F;code&gt;) files in VSCode using the &lt;code&gt;nbstata&lt;&#x2F;code&gt; kernel.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;an-example&quot;&gt;An example&lt;a class=&quot;zola-anchor&quot; href=&quot;#an-example&quot; aria-label=&quot;Anchor link for: an-example&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;To begin, create a &lt;code&gt;.qmd&lt;&#x2F;code&gt; file. At the very start of this file, incorporate a YAML header to specify the Jupyter kernel you wish to use. In our case, to use &lt;code&gt;nbstata&lt;&#x2F;code&gt; in your Quarto document, add &lt;code&gt;jupyter: nbstata&lt;&#x2F;code&gt; within the YAML header. The YAML header is a section located at the top of Markdown files, designed to hold metadata such as configuration settings and document properties. This metadata is structured in a readable format and is essential for defining various aspects of your document&#x27;s processing and presentation. A sample of YAML header could look as follows:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo z-code&quot;&gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;---&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;title: &amp;quot;Some dataviz with Stata and Quarto&amp;quot;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;date: 2024-24-01&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;author: &amp;quot;Pablo&amp;quot;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;jupyter: nbstata&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;format: &lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;  html:&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    self-contained: true&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    toc: true&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    toc-depth: 3&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    toc-float:&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;      collapsed: false&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;      smooth-scroll: true&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;      width: 20%&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;---&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;You can now create code chunks as you would do in any other notebook-based interface, specifying &lt;code&gt;stata&lt;&#x2F;code&gt; in your code chunk header.&lt;&#x2F;p&gt;
&lt;p&gt;For our example, we will create a chart to visualize gender gaps in suicide rates across low-income countries. To load the dataset into our environment, we will use the Stata wrapper for the World Bank&#x27;s WDI database API, &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jpazvd&#x2F;wbopendata&#x2F;tree&#x2F;master&quot;&gt;&lt;code&gt;wbopendata&lt;&#x2F;code&gt;&lt;&#x2F;a&gt;:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo z-code&quot;&gt;&lt;code data-lang=&quot;stata&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;`&lt;&#x2F;span&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;`&lt;&#x2F;span&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;`&lt;&#x2F;span&gt;&lt;span class=&quot;z-comment&quot;&gt;{&lt;&#x2F;span&gt;&lt;span class=&quot;z-comment&quot;&gt;stata load-data}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-entity z-name&quot;&gt;wbopendata&lt;&#x2F;span&gt;&lt;span class=&quot;z-comment&quot;&gt;,&lt;&#x2F;span&gt;&lt;span class=&quot;z-comment&quot;&gt; indicator(SH.STA.SUIC.FE.P5; SH.STA.SUIC.MA.P5) clear long&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-entity z-name&quot;&gt;desc&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;```&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;&lt;pre class=&quot;giallo z-code&quot;&gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Metadata for indicator SH.STA.SUIC.FE.P5&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;-------------------------------------------------------------------------------&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    Name: Suicide mortality rate, female (per 100,000 female population)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;-------------------------------------------------------------------------------&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    Collection: 2 World Development Indicators&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;-------------------------------------------------------------------------------&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    Description: Suicide mortality rate is the number of suicide deaths in&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    a year per 100,000 population. Crude suicide rate (not age-adjusted).&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;-------------------------------------------------------------------------------&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    Note: World Health Organization, Global Health Observatory Data&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    Repository (http:  apps.who.int ghodata ).&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;-------------------------------------------------------------------------------&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    Topic(s): 8 Health&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;-------------------------------------------------------------------------------&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    Metadata for indicator SH.STA.SUIC.MA.P5&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;-------------------------------------------------------------------------------&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    Name: Suicide mortality rate, male (per 100,000 male population)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;-------------------------------------------------------------------------------&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    Collection: 2 World Development Indicators&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;-------------------------------------------------------------------------------&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    Description: Suicide mortality rate is the number of suicide deaths in&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;...&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;sh_sta_suic_m~5 float   %8.0g                 SH.STA.SUIC.MA.P5&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;-------------------------------------------------------------------------------&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Sorted by: countrycode  year&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;     Note: Dataset has changed since last saved.&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Output is truncated. View as a scrollable element or open in a text editor. Adjust cell output settings...&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Let&#x27;s create a simple dumbbell chart by sex:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo z-code&quot;&gt;&lt;code data-lang=&quot;stata&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;`&lt;&#x2F;span&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;`&lt;&#x2F;span&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;`&lt;&#x2F;span&gt;&lt;span class=&quot;z-comment&quot;&gt;{&lt;&#x2F;span&gt;&lt;span class=&quot;z-comment&quot;&gt;stata gen-chart}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;*&lt;&#x2F;span&gt;&lt;span class=&quot;z-comment&quot;&gt; Prepare data&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  cap ren sh_sta_suic_fe_p5 fem_rate &lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  cap ren sh_sta_suic_ma_p5 male_rate &lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  keep if !inlist(.,fem_rate,male_rate) &#x2F;&#x2F; keep non-missing pairs of female and male suicide rates&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  bys countrycode (year): keep if _n == _N &#x2F;&#x2F; keep latest country-year obs.&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  drop if region == &amp;quot;&amp;quot; &#x2F;&#x2F; drop regional aggregates&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  keep if incomelevelname == &amp;quot;Low income&amp;quot; &#x2F;&#x2F; keep low-income countries&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;* Chart settings &lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  distinct countrycode&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  local ub = `r(ndistinct)&lt;&#x2F;span&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;&amp;#39;&lt;&#x2F;span&gt;&lt;span class=&quot;z-comment&quot;&gt; &#x2F;&#x2F; N of categories for the x-axis&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  sort male_rate &lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  cap g order = _n &#x2F;&#x2F; order of x-axis categories&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  labmask order, val(countryname)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  su male_rate, d&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  local x_max = round(r(max),0.1) &#x2F;&#x2F; max value for y-axis&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;* Export chart&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  #d;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;  twoway (rspike male_rate fem_rate order, lcolor(gs14%90)) &lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;    (scatter male_rate order, mcolor(midblue)) &lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;    (scatter fem_rate order, mcolor(pink)), &lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;    ytitle(&amp;quot;Suicide mortality rate (per 100,000 population)&amp;quot;) &lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;    xtitle(&amp;quot;&amp;quot;) &lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;    ylabel(0(5)`x_max&lt;&#x2F;span&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;&amp;#39;&lt;&#x2F;span&gt;&lt;span class=&quot;z-comment&quot;&gt;)&lt;&#x2F;span&gt;&lt;span class=&quot;z-comment&quot;&gt; &lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-comment&quot;&gt;    xlabel(1(1)`ub&lt;&#x2F;span&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;&amp;#39;&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span class=&quot;z-support&quot;&gt; angle&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span class=&quot;z-constant&quot;&gt;45&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;span class=&quot;z-support&quot;&gt; labsize&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt;vsmall&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;span&gt; valuelabel&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;)&lt;&#x2F;span&gt;&lt;span&gt; &lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-support&quot;&gt;    legend&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt;order&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;(&lt;&#x2F;span&gt;&lt;span class=&quot;z-constant&quot;&gt;2&lt;&#x2F;span&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt; &amp;quot;&lt;&#x2F;span&gt;&lt;span class=&quot;z-string&quot;&gt;Males&lt;&#x2F;span&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;&amp;quot;&lt;&#x2F;span&gt;&lt;span class=&quot;z-constant&quot;&gt; 3&lt;&#x2F;span&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt; &amp;quot;&lt;&#x2F;span&gt;&lt;span class=&quot;z-string&quot;&gt;Females&lt;&#x2F;span&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;&amp;quot;&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;)&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt; pos&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;(&lt;&#x2F;span&gt;&lt;span class=&quot;z-constant&quot;&gt;11&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;)&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt; row&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;(&lt;&#x2F;span&gt;&lt;span class=&quot;z-constant&quot;&gt;1&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;)&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt; size&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;(&lt;&#x2F;span&gt;&lt;span class=&quot;z-constant&quot;&gt;3&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;)&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;span&gt; &lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-support&quot;&gt;    title&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt;Gender&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt; gaps&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt; in&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt; suicide&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt; rates&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt; across&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt; low&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;-&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt;income&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt; countries&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt; pos&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;(&lt;&#x2F;span&gt;&lt;span class=&quot;z-constant&quot;&gt;11&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;)&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt; size&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;(&lt;&#x2F;span&gt;&lt;span class=&quot;z-variable z-parameter z-function&quot;&gt;medium&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;)&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-keyword&quot;&gt;  #&lt;&#x2F;span&gt;&lt;span class=&quot;z-keyword&quot;&gt;d&lt;&#x2F;span&gt;&lt;span&gt; cr&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;`&lt;&#x2F;span&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;`&lt;&#x2F;span&gt;&lt;span class=&quot;z-punctuation z-definition z-string&quot;&gt;`&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;&lt;img src=&quot;&#x2F;img&#x2F;suicide_rates.png&quot; alt=&quot;Alt text&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;h1 id=&quot;caveats&quot;&gt;Caveats&lt;a class=&quot;zola-anchor&quot; href=&quot;#caveats&quot; aria-label=&quot;Anchor link for: caveats&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;Sadly, the integration of Stata with VSCode and Quarto is far from perfect, specially if you are used to the native Stata editor and its functionalities. Here are a couple of issues that bother me:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;You will be missing the data browser and the variable explorer. This is a big one. The Stata data browser is also incredibly smooth and fast compared to the alternatives in R and Python, so it is a shame to lose one of its comparative advantages. Although &lt;code&gt;nbstata&lt;&#x2F;code&gt; comes with &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;hugetim.github.io&#x2F;nbstata&#x2F;user_guide.html#magics&quot;&gt;Jupyter magics&lt;&#x2F;a&gt; that allow for browsing, such as &lt;code&gt;%browse&lt;&#x2F;code&gt; for the main dataset and &lt;code&gt;%frbrowse&lt;&#x2F;code&gt; for frames, so far I have not figured out a way to expand the cell output in VSCode (nor in Jupyter itself), so this is not useful. If you know of a solution, please reach out!&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;You will not be able to change the font for your charts within VSCode. Not a big deal, you can always change it by opening Stata itself and changing the defaults with the &lt;code&gt;graph set&lt;&#x2F;code&gt; commands, but it is a bit annoying&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h1 id=&quot;troubleshooting&quot;&gt;Troubleshooting&lt;a class=&quot;zola-anchor&quot; href=&quot;#troubleshooting&quot; aria-label=&quot;Anchor link for: troubleshooting&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h1&gt;
&lt;p&gt;In case you are finding problems with your setup, you should verify your installations and ensure that the Jupyter kernel is correctly set up and recognnized by your system.&lt;&#x2F;p&gt;
&lt;p&gt;To check if Quarto is installed and determine its version, open your command prompt (or terminal on macOS and Linux) and enter the following command:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo z-code&quot;&gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;quarto check&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;If Quarto is not found, ensure it is installed and that its installation directory is added to your system&#x27;s PATH environment variable. Reinstalling may also help.&lt;&#x2F;p&gt;
&lt;p&gt;To verify that Python is installed on your system and to check its version, run:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo z-code&quot;&gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;python --version&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;To check your Python installation directory, run&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo z-code&quot;&gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;where python&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This command should return the version of Python installed on your system. If Python is not installed or if there are issues with the installation, this command may return an error or indicate that Python is not recognized. Similar to Quarto, ensure Python&#x27;s installation directory is in your PATH. If you have multiple Python versions installed, ensure the correct version&#x27;s path is prioritized.&lt;&#x2F;p&gt;
&lt;p&gt;To check which Jupyter kernels are installed on your system, use:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo z-code&quot;&gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;jupyter kernelspec list&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This command will list all available kernels, including their names and installation paths. If you&#x27;ve installed a kernel (such as &lt;code&gt;nbstata&lt;&#x2F;code&gt;) but it doesn&#x27;t appear in this list, there may be an issue with how the kernel was installed or registered with Jupyter.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;common-issue-quarto-does-not-recognize-the-python-installation&quot;&gt;Common issue: Quarto does not recognize the Python installation&lt;a class=&quot;zola-anchor&quot; href=&quot;#common-issue-quarto-does-not-recognize-the-python-installation&quot; aria-label=&quot;Anchor link for: common-issue-quarto-does-not-recognize-the-python-installation&quot; style=&quot;visibility: hidden;&quot;&gt;#&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;If you have problems with Quarto finding the correct Python installation, you can try the following steps (here I have assumed you have installed Anaconda, but the same logic applies to Miniconda or any other Python distribution):&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Disable app execution aliases for Python&lt;&#x2F;strong&gt;: Open Settings, go to Windows Settings &amp;gt; Apps &amp;gt; Apps &amp;amp; features, and manage execution aliases to disable Python aliases (i.e., turn off the toggle(s) that point to the Microsoft Store app, &lt;code&gt;python.exe&lt;&#x2F;code&gt; and &lt;code&gt;python3.exe&lt;&#x2F;code&gt;). By disabling these, you prevent Windows from redirecting Python commands to the Microsoft Store app, allowing your system to use the Anaconda Python executable instead.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Adjusting the PATH order&lt;&#x2F;strong&gt;: go to System Properties &amp;gt; Advanced &amp;gt; Environment Variables, then select the &quot;Path&quot; variable under User variables or System variables and click &quot;Edit&quot;. In the Edit Environment Variable window, find the entry for &lt;code&gt;C:\Users\[your-username]\AppData\Local\Microsoft\WindowsApps&lt;&#x2F;code&gt;. You can either remove this entry or move it down the list so that the paths to your Anaconda installation appear above it. This changes the order in which Windows searches these paths for executable files. To check the order of your PATH environment variable, type &lt;code&gt;where python&lt;&#x2F;code&gt; in the Anaconda prompt (usually, it should be installed in &lt;code&gt;C:\ProgramData\anaconda3&lt;&#x2F;code&gt;). Then, add this path to the PATH environment variable and move it to the top of the list, adding also the &lt;code&gt;Scripts&lt;&#x2F;code&gt; folder (usually, &lt;code&gt;C:\ProgramData\anaconda3\Scripts&lt;&#x2F;code&gt;).&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;hr &#x2F;&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;1&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;1&lt;&#x2F;sup&gt;
&lt;p&gt;Stata though has its own matrix-based programming language, &lt;em&gt;Mata&lt;&#x2F;em&gt;. Some of the most popular user-written Stata packages, like &lt;code&gt;reghdfe&lt;&#x2F;code&gt; for regressions with high dimensional fixed effects, are largely written in Mata.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;copilot&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;2&lt;&#x2F;sup&gt;
&lt;p&gt;For open-source languages, like R or Python, Copilot is a game-changer. For Stata, it is less useful, although I have to say that I am pleasantly surprised with the quality of its suggestions relative to ChatGPT 4.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;3&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;3&lt;&#x2F;sup&gt;
&lt;p&gt;For earlier Stata versions, an alternative would be using the &lt;code&gt;stata_kernel&lt;&#x2F;code&gt; Jupyter kernel. I have not tested it, but I have good references from colleagues. You can find the installation guide &lt;a rel=&quot;nofollow noreferrer external&quot; href=&quot;https:&#x2F;&#x2F;kylebarron.dev&#x2F;stata_kernel&#x2F;&quot;&gt;here&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;4&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;4&lt;&#x2F;sup&gt;
&lt;p&gt;&lt;code&gt;nbstata&lt;&#x2F;code&gt; is a Jupyter kernel. If you install Anaconda, you will also get JupyterLab installed, so you will be able to run Stata code from Jupyter notebooks as well. I personally hate them, but if you don&#x27;t, it is an alternative.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
</content>
	</entry>
</feed>