Speaker: Adel Daoud
–Transcript–
Organizer: Adel and I started talking to each other, and I began to learn about his work on global poverty, which you will see a great deal of today. However, the basic idea—without stealing his thunder—is that he was essentially using deep learning algorithms to connect satellite image data with ground survey data. He will discuss this in much greater detail. At the outset of our collaboration, he asked how we could recruit more students for this ambitious project. I replied that I had a proven approach through the Astrostatistics group, where we organize workshops. Indeed, this group functions primarily as a workshop series. Thus, he has been familiar with Astrostatistics from the very beginning. Initially, we considered establishing a separate Astrostatistics for this initiative, but our discussions shifted toward methodology, as I needed to develop a new course now in its third year—course 288. That marked our starting point. Together, we designed this course, which some of you have heard me describe as emphasizing multi-resolution analysis through multi-source, multi-faceted perspectives. You will observe today in Adel’s presentation how these elements align perfectly. Consequently, it was an ideal synergy: I was advancing theoretical frameworks and suddenly encountered this compelling real-world global challenge.
That is how our two endeavors converged, yet we always aimed to link with your group, as Adel’s work undoubtedly incorporates technologies and mathematical tools that could prove valuable here. In essence, he seeks to employ deep learning models and convolutional neural networks—whose mechanisms remain somewhat mysterious—to integrate data from diverse sources for predictive purposes and deeper insights. Now, correct me if I am mistaken, but he inquired whether to address causal inference, and I advised that astronomers likely contemplate it less frequently than social scientists. It is not that you avoid causality; you engage with it extensively, supported by robust theories, often tackling complex problems by scrutinizing data to derive meaning. Am I accurate? Even prominent causal inference experts suggest it is not overly prominent. Alright, noticing the silence, I appreciate it. Essentially, Adel’s objective is to construct a database akin to a softened version of physics, much like how Connor here conducts extensive causal inference. Ultimately, our efforts pursue two aims: creating databases for broader use, as you do, and discerning how global poverty indicators influence policy and human lives, thereby incorporating the causal dimension. I invited him to speak here, acknowledging the challenge that his research is Earth-bound, yet he devised an elegant bridge to your domain. Let us assess its efficacy.
Adel Daoud: Thank you for that fantastic introduction, J. The pleasure is mine to collaborate with you. My challenge for this talk was to forge a connection between our lab’s work and your impressive endeavors, which we admire in the news and elsewhere. As outlined in the abstract you have read, I framed this through the lens of life in the universe. We know of one planet harboring life: our own, evidently. Accordingly, what we have been doing here in this lab is really to explore indicators of life and well-being through advanced data analysis techniques.
Adel Daoud: Let us consider ways of measuring living conditions. Is this okay? Yeah, let me remove this. Wait, no—hit “more,” okay, then go down to “hide.” Sorry about the floating meeting controls. All right, is it good now? Yeah. If there are any questions online, please help me with those, as I doubt we’ll be able to see them. I expect numerous questions along the way. Thus, the manner in which I envision this talk unfolding is to reflect on life in the universe and on our planet, exploring ways to generalize insights from our work with the lab, alongside James, Connor, and beyond. Are there any overlaps? Obviously, we are still awaiting the moment when ET will appear; at least, I harbor hope each morning while reading the news for a breakthrough on that front. However, since we continue to wait, the assumption I will adopt in this presentation is to discuss ET as a fictional character. In general, I will address life on Planet X shortly, within the structure of this talk. My mode of argument posits that lessons from measuring life on this planet might, to some extent, generalize to assessing life on other planets; the nature of these measurements will vary based on the types of data accessible. This forms the parallel I have in mind. Again, I am neither an astrophysicist nor an astrobiologist—perhaps we are witnessing the inception of astrosociology. Why not? I initially thought “astrodemography” might be a more suitable term, but upon searching, I discovered it was already taken. Wikipedia revealed the term “astrosociology,” which I had never encountered; I believed I was pioneering it, yet here I am, twenty years late—it seems to have originated around 2003, when someone began developing it. Thank you; I am learning myself. You might immediately inquire how we can tackle such a daring question. As Xi explained at the outset, we have been addressing these issues for quite some time now. Although I serve as the lead PI for the AI and Global Development Lab, I truly view it as a collaboration—as I mentioned yesterday—with multiple PIs and independent researchers uniting to forge something innovative. The virtual lab we lead constitutes a consortium of researchers here at Harvard, with James Ji (and perhaps others will join), Connor (now at Harvard from Texas), and my collaborators, PhD students, and postdocs in Sweden at the computer science departments of Chalmers Technical University, as well as the Institute for Analytical Sociology at Linköping University. We examine questions in the domain of global development: what is occurring on the planet, and how can we deploy tools to measure planetary-scale processes crucial for sustainable development, particularly focusing on poverty and health. Our interests center around three areas: measurement, causal inference, and their integration for applied research. This picture illustrates the team spanning disciplinary boundaries, including statisticians, computer scientists (as mentioned), social scientists, economists, and political scientists. While awaiting ET’s arrival, my goal with this talk is to maximize the transfer of knowledge from what we have learned, and how that might prove helpful in the future for measurements.
Adel Daoud: Here on Earth, as on other planets, we have unequal conditions, so we are pretty confident in at least what we can say for this case. The structure is that I am going to discuss measurement endeavors we are undertaking in the continent of Africa, along with a bit about India; these examples will demonstrate various methods and modes for measuring different outcomes of interest. I will monitor my time allocation, thus you must help keep track of me, as I anticipate conversations arising along the way. I plan to emphasize the Africa results more heavily, but if time permits a deeper exploration of the India findings, that could prove useful. I also encourage you to interrupt me as needed; Emil highlighted that there would be numerous stops, so let us proceed accordingly. Rather than delving deeply into the mechanics of the methods—though I am happy to address them—I believe it is more valuable to concentrate on their potential applications, yet certain technical aspects are crucial, so feel free to inquire. Toward the conclusion, I will present two or three more substantial slides to consider what it would entail to generalize our measurements to other planets, including the requisite data types and measurement approaches; here, I truly hope you will collaborate in pondering these issues.
While flying from California, where I am currently at Stanford’s Center for Advanced Study, I was reading an unrelated book, yet my thoughts were brewing on this presentation, leading me to generate exciting ideas for conducting such measurements on other planets—and at one point, I considered whether a commentary piece might be warranted on this topic. Therefore, if anyone is interested, perhaps we can discuss it later; it is not a fully developed research article, but rather some preliminary thoughts and lessons learned, so inform me if those insights align with that line of work. Alright, thus let us concentrate on Earth and the ongoing developments there. One critical activity in our lab involves addressing data gaps, because studying poverty, health, and living conditions across the planet in general—not solely in Africa, though we prioritize it for practical reasons—requires abundant data.
What you see here is an illustration of diverse living conditions observed globally: on one side, access to drinking water varies in quality across humanity; modes of transportation differ widely, with many relying on foot travel as our ancestors did for hundreds of thousands of years before inventions like bicycles emerged, or before we tamed horses and developed such innovations. Living conditions encompass multiple dimensions, including access to water and transportation, as well as cooking methods. I would argue that the second option is actually the healthiest, progressing to the least healthy. That is an excellent observation, as there is certainly a historical shift toward cars, followed by a return to cycling—and whether that is beneficial or detrimental remains debatable, even extending to walking; indeed, I notice James often opts for the stairs instead of elevators, which benefits my body at least. Regarding data, we recognize that the distribution of resources among humanity is vastly unequal, with many people living in low levels.
Level four includes us having food, water, housing, shelter, and transportation provided in the highest quality way that we can experience. Obviously, there are variations within rich countries, but on average, we can think about people in the industrialized world living in level four conditions. Level one is what concerns most policymakers when we think about poverty and health, whereas level four is where we also consider the environmental impacts of our lifestyle. However, we are not focusing on environmental effects in this talk; instead, the critical aspect that global development researchers like myself and others are thinking about is how to push more and more people up the level ladder.
Audience Member: Right. I have seen this picture many times, but I actually never noticed what the dollar signs there mean.
Adel Daoud: Yes, that represents the World Bank threshold of two dollars per day on average. Thus, these are approximate thresholds of what it means to be poor or living in levels one, two, and three, which form the middle range. These numbers are contested from a monetary perspective or in terms of monetary definitions, in the sense that the World Bank focuses very much on income, whereas there has been a growth in thinking about the consumption side—what do we actually have access to? Therefore, two dollars per day does not reveal much in terms of what food I consume or what transportation I use.
Audience Member: Are there also quantiles, I mean, in some like the 25th percentile?
Adel Daoud: Yes, there are definitely ways of making this graph more fine-tuned toward the distribution of living conditions. Ultimately, to achieve that—and really what we want to do—there is a lot of income data if you look at the global scale, though there is not as much in Africa still, as I will show in a couple of slides. Hence, what we really need is data at high resolution, meaning that we want it to be temporally and geographically extensive in coverage.
Audience Member: Perhaps a related question: When the UN decides, do they say the lowest is 10 percent or some percentile?
Adel Daoud: Indeed, it is interesting; the definition itself is more based on an absolute measure of what would be needed. Although it is still arbitrary, it is basically a consumption basket that they have defined, so to satisfy this consumption basket and be able to buy what is required on average, they come up with this measure. But if you consider a totally different context, such as the US, where we think about what it means to be poor there, we then consider 60 percent of the median income—that is a rough measure used based on the distribution of incomes in the US; if you are below that, you are then considered poor in the US.
Audience Member: Let me interject, sorry—please. That actually, so the world learned this from David Gordon, right? He was talking about one challenge here: it used to be, because I do a lot of surveys, that when talking about surveys translated into different languages, you want to make sure the same thing translates to the same level so everybody understands. In this case, poverty is actually the other way around, because he was talking about countries like this where the question, for example, “Do you have access to the internet?” could indicate you are poor if you do not have it. But in other countries, do you have a habit? Thus, you actually want the questions to work very differently by trying to capture the social question.
Adel Daoud: The design question is quite interesting: it is not about maintaining consistency but preserving the meaning, which is much harder. In essence, these dimensions offer latent factors for capturing such aspects, yet the methods of measurement vary across the world. The challenge for those of us seeking to create harmonized data for the planet lies in integrating this information to consider the human species as a whole. Thus, how do we approach that dataset from this perspective? In our particular focus on Africa, we regard these outcomes as the essential data required; regardless of how one analyzes it, a disaggregated labeled dataset on living conditions across the planet remains necessary. The issue we are addressing, as alluded to earlier, stems from data scarcity in many locations on Earth—both spatially and historically—prompting our aim to develop tools and methods to bridge those gaps.
Currently, we are demonstrating some existing variation using satellite imagery. Therefore, if surveys are insufficient, how might we integrate them with satellite images? These are Landsat images from 1984 in Cape Town, illustrating temporal changes and visual variations that reveal human activities on the planet. The key question involves combining this data with known living conditions through computer science and statistical methods to enhance measurements of these conditions, with temporality being especially crucial. For instance, focusing on Cairo in Egypt, a cross-sectional dataset shows little change, but examining the period from 1984 onward in a specific neighborhood reveals that nothing existed there initially, only to be built up over time—indicating human development in that area.
Consequently, a critical project goal is to address data gaps, as seen in countries like Burundi where certain time slices lack information, though satellite images are nearly always available. We aim to train deep learning functions that take images as input to estimate living conditions, yielding a dataset like the one discussed earlier to fill those voids. This example is merely one case; we target the entire continent across that timeframe, where without our project, only the 2010 and 2017 Demographic Health Surveys would be usable. These surveys, conducted in Africa since the late 1980s and commissioned by the USA, require enrollment from aid-receiving countries. Once trained, the algorithm enables gap-filling, providing time-series data on ground-level developments. To specify the training process further, we employ a supervised method necessitating ground truth data, and so the points here that you see represent that.
These are measured locations of those dimensions that I showed you earlier. Thus, we create this gridded map displaying living conditions in the locality. The goal is to generate the heat map on the right, where we divide these geographical points, training on the blue ones and then imputing, predicting, or estimating living conditions in the red points, producing various informational graphs that illustrate the dynamics.
Audience Member: How do you identify these field sites initially—the blue ones?
Adel Daoud: The blue plots derive from actual surveys conducted on the ground, providing some data—exactly. This is a supervised method, requiring such input. As we will discuss later, generalizing to other planets would obviously be challenging. However, it emerges that these surveyed labeled data can involve either people visiting those localities to measure them or leveraging prior information about the image to label them manually without physically going there. That is how the supervised approach functions: the more information available—though labels are essential—the better the performance, yet labeling is constrained by the knowledge of that image and locality. Fortunately, in our setting, we possess numerous DHS surveys across the entire Global South. Focusing on Africa, the plan is to cover at least 60 to 70% of the global population. Additionally, consider combining data sources, which I will address in the India section of the presentation, where we integrate census data, survey data, and perhaps other types to triangulate these estimates.
Audience Member: So the other question is, if you don’t have this known input, you cannot generate the new locations of interest. You mean the satellite imagery now? If you start with 2009–2011, you know these are the blue regions of interest, so do you include any new regions next?
Adel Daoud: I see what you mean. Essentially, as you noted, for 2009 we have training points in that locality, but we also possess training points from other time slices. Therefore, even for the final results, certain assumptions apply here: one is that we capture the time trend in living conditions through the data available in those temporal slices. Otherwise, if relying solely on 2009 data initially and extrapolating forward, we assume stability in the observations from that time slice persists over time. In contrast, with multiple survey points, even amid distributional changes, we presume to capture them via those new surveys.
Audience Member: What is your spatial resolution?
Adel Daoud: The spatial resolution for the Landsat imagery is 30 by 30 meters—down to a house, essentially a neighborhood. Thus, it equates to 30 meters by 30 meters per pixel. We employ Landsat imagery, which I will cover in a few slides; it offers historical coverage and is publicly available. However, newer satellites launched in the past decade provide sub-meter resolution, enabling identification of even the car type parked in a lot.
The question was about the resolution of the poverty map, and that is exactly right. There are several layers to that question; one is that the actual surveying is conducted face-to-face in a household, with a DHS surveyor going out, knocking on the door, and administering a questionnaire inquiring about the living conditions of that particular household. In a neighborhood, these points represent households and are collected within 5km radii; the surveyors gather GIS coordinates over the centroid of these points, aggregating the numbers as a means of protecting privacy.
Furthermore, in the paper we have been working on with Jal and James, along with a collaborator, the GIS point is displaced by a random angle and random distance to another location to safeguard privacy, thereby injecting noise and diminishing the precision of our estimates. In this initial paper, we utilize prior information combined with multiple imputation to mitigate that bias.
This illustration depicts the availability of Landsat images; Landsat has existed since the 1970s, with the technology evolving over time, yet from Landsat 5 onward, the resolution has remained constant, although additional bands have been incorporated to measure various environmentally related activities on the planet. Our work primarily emphasizes the RGB bands to visually document events on the planet, and we integrate this with night light imagery to observe nocturnal activities. Considerable research over the past two or three decades has employed night light data as a proxy for economic development; accordingly, the concept is that highly developed areas like Boston or New York are brightly illuminated.
Audience Member: Maybe that was a muting issue. How much do we know about the improvements to satellites over time, beyond the resolution increasing? Are other aspects being enhanced, such as coverage?
Adel Daoud: For Landsat, there has typically been one satellite orbiting the Earth at a time, whereas the European Space Agency launched Sentinel in 2014 with two satellites orbiting simultaneously. Consequently, the revisiting rate is approximately 16 days—or two weeks—for each locality with Landsat, while Sentinel enables coverage every three days on average. With commercial satellites, the objective for some is to conduct multiple measurements; indeed, the aim is to perform within-day observations, potentially in the morning, midday, and afternoon. Certain systems feature Dove satellites, which are roughly this size, with about 200 orbiting to establish a network for temporal measurements. These represent truly compelling products suitable for contemporary assessments, although our focus encompasses historical global development as well; nevertheless, looking ahead, our project certainly holds potential for incorporating such imagery.
Adel Daoud: I have a question for the Astro friends: do you guys use these kinds of datasets at all? Is this completely useless to you?
Unknown Speaker: No, you should be very familiar with the GOES, which I don’t know. GOES, right? I mean, this is the Pleiades data that we’ve been…
Adel Daoud: Looking at this, it is a different type of data. We do not have the same type of repeated observations. That is right. However, this particular one is the new satellite TEMPO, which is looking at the atmosphere and monitoring pollution in the area. It was just launched last year, and there is the Center for Astrophysics to analyze this data. Therefore, this is a new type of data, right?
Audience Member: What is the setup of that data collection?
Adel Daoud: Oh, I am sorry, but mainly it is atmospheric and involves a combination of a few satellites. There was a Japanese version launched earlier, and this is the American version. It is all in collaboration; they monitor the entire Earth and look at pollution. I do not know the details about some specific elements, so I have to learn more about that. All right, that was a bit of the data background. Let me dive into some of the results that we have for Africa.
This is mainly based on a paper that my PhD student Markus Pettersson led and published in the proceedings of the International Joint Conference on Artificial Intelligence, special track on AI for social good. This is one of the early data products or methods that we have created and published last year. The main innovation there is really the time series aspect; a lot of these predictions have been floating around but focusing mainly on cross-sectional data, and I will be talking a little bit about how we achieved that sort of temporal architecture to capture that type of information. As I said earlier, Africa is the focus.
We have about 60,000 DHS survey points; these are cluster points, which you can basically see as neighborhood-level measurements as released to us, but we have household-level information. Let me rephrase that: the neighborhood-level GIS coordinate is what we have, but we have household-level information about these outcomes that I talked about earlier. There are about 36 countries, and the time slice is about 30 years. The goal is to see how well we can build this deep learning model to create the dataset of poverty levels.
Audience Member: So, do you see the data as a map or as point data?
Adel Daoud: Yes, I mean, we see it as the end result that we have created, which is 30 by 30 meter resolution. That is sort of what comes out of the model, but we could technically create a point estimate by basically creating a moving window of predictions at an even finer scale if needed or wanted. However, the question remains the quality of those predictions, which we have not evaluated yet.
Audience Member: It is more like, you know, when I say map versus points, a map is thematic, right? Even though you may not have an observation at a particular location, you can imagine that there exists an intensity that you observe. Whereas with point processes, there is a point measurement at one location and nothing else.
Adel Daoud: Yes, yes, I would say the outcomes of the model are points, but our goal is maps, if that makes sense.
Audience Member: Actually, that is a really good point. No, it is an interesting point because, in the old days, before the deep learning boom, we always thought about modeling as a spatial process. But now, these methods simply treat them as pixel values, though effective; actually, it captures things that should be there because the spatial patterns obviously should be present. I guess the actual implementation you used puts this CNN such that it actually has pixel values for every location, right? Because even if the site is missing, it will give you some value there.
Adel Daoud: That is right; that is the interesting part of the modeling process, that these…
Adel Daoud: The blue points that you asked me about are basically the ones we are showing on this map right now. These represent the locations for which training data are available. Essentially, to provide more intuition about the input and output, you are limited by the resolution of the Landsat image, which consists of 30 by 30 meter pixels. The model requires this dataset, comprising the images at that resolution. Those points sit at the center of each image, and we have a scalar value of y that the model aims to produce. What the model does is find the mapping between the image space and that point space. Although the output is at points, nothing prevents you, once the model is trained, from scanning over new points in whatever range or window you desire to approximate a map, if that makes sense. In our predictions, due to computational constraints as well as practical limitations, we generate predictions only where people are known to live. Thus, we employ another human settlement map as a cookie cutter overlaid on the African continent to restrict predictions to those areas. Furthermore, to discuss the outcome variable more precisely, we use what we call the International Wealth Index, an asset index derived from the Demographic Health Survey. Each column here corresponds to a questionnaire item in the survey. For instance, if a household reports having a television but no refrigerator, no phone, and similar responses, it would receive an IWI value of 12.73, placing it somewhere on a scale where 100 is the maximum and zero represents absolute poverty. You would fall into this part of the distribution with such answers, just to offer some intuition.
Audience Member: I’m curious, what does “utensil” refer to here?
Adel Daoud: They distinguish between cheap and expensive utensils, essentially based on the quality of the kitchen utensils you possess. Therefore, in our case, we would probably not qualify as having expensive ones at this particular moment. That’s it. The outcome variable ranges between 0 and 100, constituting a principal component analysis measure over these indicators. I aimed to convey what zero and 100 signify, along with the impact of changes; for example, without a television, the score would drop from 12.73 to 4, positioning the household at the absolute bottom.
Audience Member: Interesting. For instance, if I lack a TV these days—most of us are eliminating traditional TV watching, though we still consume content— it highlights the cultural issues with these questions.
Adel Daoud: It is challenging to devise such measurements even within a single country, and even more so when considering temporal aspects alongside the cultural shifts we observe. Consequently, conceptualizing utensils for other planets like Orion would be difficult, yet this is the reality we must address on Earth. Hence, as a key takeaway, what you need for that is recognizing certain bottom questions as more reliable, such as access to drinking water and floor material. I agree that those are very dependable. The way I view these items is as indicators; although they serve as direct measures, in essence they reflect underlying socioeconomic conditions.
Adel Daoud: We are trying to indirectly capture the living conditions in the household. Although they might be inaccurate for some households, on average, the data will be able to capture living conditions.
Audience Member: The number of rooms also depends on how many people there are in the household.
Adel Daoud: That’s exactly right, and that will also vary. Well, there is Jeff’s comment in the chat: in case you don’t know about the NASA data visualization facility at Ames close to SFO, the Hyperwall may be used for your research. Hyperwall, Hyperwall. Yes, if you can please put that link in a separate text file, I’ll look at it later. Thank you very much for that comment. Alright, therefore we are building intuition here about the outcome and what we do on this planet; moreover, if you want to do similar maps on other planets, you would likely need to follow a comparable approach.
The satellite image data is again the critical piece. Of the bands that we are relying on, these are the RGB bands; we have other bands sitting there in the Landsat imagery that we’re working with, but also the night light bands that we add to our data, which do not exist in Landsat and are measured by different satellites. In this particular paper, to handle problems of clouds—which is a recurring problem in Earth observations and that is also a common problem in our work—you have clouds on planets, yes. Consequently, I figured clouds are becoming one of the more evil things that I’m thinking about; I’m having nightmares about clouds, not only because they create bad weather but also because they interfere with our data measurements as well.
We need to create an anti-cloud party or something like that. I would say the three-year median is what we did in this paper; it’s something that we are improving, but I feel this is a bit arbitrary of why three years. There are seasonalities that are happening within the year that are perhaps important to capture, especially if you think about agricultural development or even fast-moving processes that happen on the planet, which could be natural disasters or conflict. If you want to create data like this, you want to make sure the temporal variation is as high as possible as well. Hence, what we want to do here is to capture slow and fast processes; right now, we are stuck in this paper with capturing more slow-moving processes.
Thus, what we’re doing here is then to slice up the temporal data into chunks of three years, so we call these frames; what we then do is create these deep learning models that are able to take frames as data. The state-of-the-art has been doing something like this, which is that it takes one single frame at a time; what we did, and this is the big contribution of this particular paper, is that we added what we call an LSTM layer—a long short-term memory layer—where you can link information that sort of comes down from the upstream of the model, and so the information of what happened in the past can propagate forward in time. That’s a fairly standard way of evolving things you might say, but it’s actually making quite a difference.
Audience Member: Yes, I have a question related to these clouds and your pixel data points only. Can you use also the weather information for the areas so it will give you a better resolution than just ten images?
Adel Daoud: Yeah, we do; we do have, because of the approximate 16-day revisiting rates, technically we would have about 24 images per year.
We average images per locality within a year for two reasons: first, to tackle clouds; second, to handle the massive amount of data available to us. Therefore, this is a pragmatic decision that I would improve upon today. Indeed, we are now including that information to the extent possible. One might ask whether we need all 24 images, as that would significantly increase the computational demand. Could we use one representative image per season or per month? These choices are to some extent substantively driven, but also pragmatic, which is characteristic of the science.
Audience Member: Please, may I ask: You mentioned that you have this population presence-absence mask that you’re applying, so you can figure out which pixels you’re actually going to train on, because they have people there and therefore you might have survey data. As people expand further out from cities—we saw this in the Kenya example at the beginning—are you changing this mask over time? If so, what is this doing to your data dimensions when you go to train on the data in this time-series manner?
Adel Daoud: That’s a really good question. The human settlement mask is applied at the post-training, post-prediction stage, after we have trained the model while confining ourselves to the DHS data by necessity; we wish we had more data points, so we use whatever is available. Then, we make truly out-of-sample predictions, meaning we have no labels and are trusting the model to create the data. At that stage, we apply the cookie cutter, focusing only on human settlements, because we don’t want to apply it to the whole continent, as that’s computationally demanding. It’s also nonsensical: if you predict in the Sahara Desert, you’ll get a prediction that is probably not zero, and it will not be exactly zero. I’ll show you some of that, which gives you a sense that the model is always like an autistic model; it’s just doing what it’s trained for, predicting between zero and 100 by trying to fit any given image to that prediction.
Audience Member: When you train, do you have zeros going in for the Sahara Desert?
Adel Daoud: We have not, but actually, it’s a really interesting idea. Perhaps we need another class, since we are confining it to 0-100, assuming that everything fits that universe, but it does not. Therefore, if you want a more realistic model, you would have to add another possibility, such as identifying non-human settlements and assigning zero accordingly. This enters the realm of semantic segmentation: if you want to predict deserts, forests, and so on, you need labels for those as well during training. Nevertheless, it’s a really important question, as it highlights the limitations of training.
Audience Member: Just one thought about bringing this to the astrophysical realm: it would be interesting to apply the model as trained in other planetary contexts, just to see what the model spits out. That could teach us something about what it pays attention to and what’s not paying attention to.
Adel Daoud: Yes, one of the lessons learned that I’ll be coming to is this: suppose we discover populations on Planet X—ET discovered, whoa!—but we want to know how they are doing. If we have no data at all, could we apply a model like this to it? You’ll see that yes, if some conditions are met. Obviously, if they are exactly like us—if they use similar utensils—we could generalize it to them.
Adel Daoud: Regarding that state, I’ll come back to it. Nevertheless, it’s an interesting thought concerning the limitations and what would be required to capture light in order to measure life or living conditions on other planets. Alright, let me know if I’m going too slow or if I’m doing well on time.
Audience Member: What is the typical size of the overall dataset or region that you look at? Do you always examine the data globally, or just in specific regions? I presume that is very interesting, since there must be regional correlations that can even extend far. For example, the wealth index in Europe and the US is high. I was wondering about the typical scale of those correlations.
Adel Daoud: Yes, that’s a really important point. As we validate the results, you’ll see much of that geospatial correlation emerging; countries closer to each other in terms of how the IWI is generated will, as you noted, perform similarly on the metric. Much of the probing in training these models—which is challenging given their large scope and size—involves examining where the errors occur, where we perform poorly, and where we excel. As you will see with Egypt, despite abundant training data, we perform quite poorly there; interestingly, many data points, both rich and poor on the IWI scale, are concentrated in the same dense area along the Nile, making it difficult for an algorithm like this to distinguish them, since the 6.7 square kilometer image size we use is probably too large for such differentiation in Egypt. This raises a question that Jolie, James, Connor, and I have pondered extensively: what is the proper resolution for the problem at hand? Drawing intuition from predicting whether an image shows a cat or a dog, the ideal resolution centers the cat or dog in the image, reflecting how the model’s attention functions; however, in a problem like this where households are scattered—like a cat in the background—it becomes harder for the model to predict. The same applies here, as it remains unclear how attention should focus, complicating the task. Ideally, to advance this method, we would use household-level pictures featuring just one household—not a neighborhood, unlike the South Africa visualization you saw—but each with its particular estimate, thereby separating out variation in the data.
Audience Member: So, if you can wrap up in about 15 to 20 minutes.
Adel Daoud: Yes, in 15 minutes. Okay, so let me move on a bit to assess our progress. This returns to the question of generalization: how well do we generalize across different training modes? What we see here involves out-of-area, out-of-country, and out-of-time span scenarios, which I’ll explain shortly; but first, consider an intuition on how machine learning models are ideally trained, since they rely on IID data.
Adel Daoud: Right, you have data that you train on, but you generalize to new data that you assume comes from the same distribution distributionally, though the model has not seen it. Thus, you’re learning and generalizing in a way that is dissimilar. Here, we have issues of geographical and temporal autocorrelation, meaning that clear separation does not fully hold.
Therefore, we train it in different ways to at least stress test the model on how it would generalize if the generalization set is defined differently. Out-of-country and out-of-time-span are relatively self-explanatory: basically, we hold out, say, Egypt, train on the rest, and predict on Egypt—but there are actually several countries there. Similarly for time span, we train on a certain time period and hold out others as such. Out-of-area just means that there’s a boundary within which DHS locations cannot overlap each other in nodes.
And this gives you a sense of the general performance, how much the frame… these are the same model, just that it takes different amounts of frames and simultaneously the train set. But we achieve a performance—this is the Pearson correlation metric—of 0.76, which, if you stop and think about it, is pretty good. Who would say that an image can give you information about utensils, stoves, and TVs and whatnot?
But this is the Pearson correlation, which has a problem where you could have a constant shift, and it doesn’t really detect it. We also computed the R-squared values; they’re in the vicinity of similar numbers, between 0.75 and up to touching 0.8 R-squared. Here, I think we also corrected for the sample mean in the data. In the interest of time, we have these plots showing a little bit of how well we are doing, where we are doing best, and where worse; breaking it down by rural and urban, you see that this is the predicted IWI and then the overall wealth that we’re predicting. The rural areas are actually where our model is performing better, because it’s easier to separate out forest areas and human settlement areas. You have the by-country results, and you can also break it down by the distribution of the IWI to see how you’re doing.
By country results—what does zero mean? Exactly: here, in terms of R-squared, you can even get negative R-squared, so basically you’re not doing better than the sample mean prediction. This is similar here: if you have zero, you’re not doing any better than just guessing the sample mean on average for that sample. As I mentioned, Egypt is pretty much in this vicinity of the predictions.
Similarly, you can break it down where we are doing better in terms of the mean error improvement, and we see that where this particular setup is working better is in the tail end of the upper part of the distribution of IWI; that’s where the predictions are better. Additionally, some are talking about deserts, and now we’re also thinking about the mean squared error: if we just visually and qualitatively probe where there are fewer mean squared error improvements, it turns out that if you select the areas of interest—I don’t see any human settlements here—this is most likely what has happened: the DHS privacy protection mechanism has displaced the DHS point somewhere into the desert or a mountain, and there’s not much information for the model to work on here. It’s basically when it gets black, there’s most likely missing data.
Regarding Landsat imagery, to build intuition, the most improved areas exhibit actual changes; for instance, this small village—though a bit hard to discern—shows developments over time. This example is from Namibia, following one from Morocco, and similarly, we focus on urban settlements in South Africa and Egypt, which in this case are progressing well: nothing existed there in 1990, but structures have since been built. Thus, validating the predictions, one might wonder what happens if we employ entirely different data that still retains information on the IWI score.
Examining the Human Development Index—a UN measure aggregated at the country level—we aggregate our numbers by country-years and assess how well our predictions align with these country-level estimates, yielding a correlation metric of 0.5, which is reasonable though not perfect. For me, this figure gains meaning only when compared to others’ achievements; in previous slides, I presented benchmarks from various models, where the single-frame approach represented the state-of-the-art before our multi-frame model, which outperforms it—they achieved around 0.73-ish, while we advanced it to 0.76, akin to a competitive race.
Moreover, we excel particularly on temporal predictions where theirs are notably weak. From this, we can generate data such as these point maps—to address your earlier question—that we have created by using human settlements to precisely extract locations of interest. With these data, as shown on the right, we conduct forecasting exercises to determine which villages might reach a certain IWI threshold by a specific year, alluding to the policy implications of how many will escape poverty; our results using the created data fall within this range here.
I have some findings from India, but I will skip them entirely and proceed to the lessons learned, which should be more engaging to discuss. To provide a broad overview, our takeaways mirror what we did in the Africa model, but the challenge in the India setup involves differing ground truths: we have surveys from the DHS, census data, and mismatched labels. Thus, the question is how to train on one set and generalize to another using transfer learning techniques, which allow information transfer from one distribution to another under certain assumptions—I will reserve details for later and simply display them here for viewing.
We also disaggregate the IWI into various measures; we examine about 90 different ones to evaluate generalization across these outcomes. Some of these, as shown in this R-squared plot from the top of the 90-list models, include everything from predicting whether women are overweight from satellite imagery, which sounds remarkable, but you must consider it as capturing a latent factor through other processes visible in the image that resemble a poor area and thus correlate accordingly.
Adel Daoud: I would love for you to sit and meditate over this table, but since it’s in the Zoom recording, you can do it later and we can discuss it then. However, I also want to show you where we perform really poorly—for instance, whether children aged 12 to 13 have received the measles vaccine, which is a question in the DHS, and we get negative R-squared values there as well.
Additionally, we generate data like this, and one interesting aspect is examining activation maps: essentially, after training the model, the question arises as to what elements prompt it to generate a prediction. For the outcome of a permanent house, this should be reversed—blue actually indicates warmer areas, though it should be red—but you can see what is activated here. When we examine the RGB bands, it’s evident that human settlements activate how the model responds to incoming information.
Looking at transport and motorization, there is a road cutting through the image, and it overlaps nicely with the road in the image itself. Thus, these activation maps are helpful for probing the model to understand what it is actually reacting to. I think this marks the concluding part, which focuses on lessons learned: what did we glean from this? Fundamentally, we discovered that we can measure living conditions through a combination of satellite imagery and training data—that’s the major takeaway.
Therefore, if you consider applying something similar in the astro-demography discipline that I’m envisioning, what conditions must we satisfy? Obviously, the first approach, which is extremely challenging and would be impossible if ET existed, involves traveling there to inquire about their living conditions—studying that and assessing how well we can conduct a DHS survey or equivalent, asking about utensils they possess, whether they have TVs, and similar details. Yet this is infeasible due to travel requirements, so instead we resort to combining planetary observations—what I might term in lieu of Earth observations—and the machine learning methods presented here. What assumptions should we consider in this domain?
Another key insight is that ET must inhabit the planet’s surface, as any subsurface dwellers would be very difficult to detect. At minimum, ET needs to leave a systematic footprint informative enough to be captured by a satellite image. As observed, humans require housing for shelter, construct transportation roads and the like, and cultivate land, thereby creating surface changes on the planet; it is crucial to systematically capture those traces of living conditions. Indeed, one implicit yet critical factor is how ET organizes society: do they reside in cities, or are they more secluded?
Consequently, even if numerous on the planet, they might inhabit very rural areas, which would influence the model. Moreover, we require technology capable of measuring the planet’s surface—for Planet X, we work with 30-meter square estimations, while the older NightLight imagery offers about 500-meter squared resolution, yet it still provides information on economic activity. If such capabilities are feasible, then we have an opportunity to employ methods like this. I conducted some research but could not identify the best resolution for surface capture.
Capabilities of satellites currently—and even possibly in the future—are limited by physical constraints. Before moving on to the next slide, I wonder what hopes there are for developing technology that could achieve 30-meter resolution, or perhaps 100-meter or 500-meter square resolution.
Audience Member: Higher resolutions are better.
Adel Daoud: Say again?
Audience Member: Milliarcsecond resolutions are better, okay.
Adel Daoud: You cannot even detect planets yet. Right now, everything we have treats planets as point sources with 0.1 arcsecond resolution.
Audience Member: That’s what I was just thinking about.
Adel Daoud: Thus, if you were to observe the Earth under such conditions, you would see developed and undeveloped patches, with your signal being a combination of both. Determining the presence of these two components would represent major progress.
Audience Member: Okay.
Adel Daoud: These satellites are orbiting the Earth and looking outwards, right?
Audience Member: Yeah.
Adel Daoud: Precisely; hence, that’s what we can work with in terms of cost and technology. Alternatively, we could send something closer to the planet of interest, but that would be costly, difficult to manage, and take a long time.
Audience Member: Take a long time, exactly.
Adel Daoud: Therefore, this approach represents a compromise regarding how close we can get to the planets of interest to capture data. Another possibility is tracking variations over time; if there’s extraterrestrial life on a planet we’re observing for 20 years, we could detect changes.
Audience Member: Yeah, yeah, yeah.
Adel Daoud: To provide context, if we sent a satellite orbiting planet X, the data size we’re dealing with currently is about 15 gigabytes for African satellite imagery and less than one gigabyte for survey data from 60,000 households. This illustrates the information that would need to be transmitted back to Earth in a similar setup, including 30-meter square data across seven bands. We could compromise a bit to indulge some fantasies. As you shake your head, I know this isn’t sci-fi—this is the fun part of the talk. Nevertheless, we still need images for predictions, although there’s some flexibility with labeled data. That’s what I’m aiming at here; the images are essential as model inputs.
Since you work with satellite images, I hope to inspire you to pursue this at some point, assuming such opportunities exist. To replicate our discussion exactly, you’d need all that data. With transfer learning, however, you could train a large model on Earth that effectively measures living conditions here and transfer it elsewhere. For instance, in the India example I didn’t fully discuss, a smaller labeled dataset suffices for fine-tuning to capture living conditions on that planet. Nonetheless, this is a highly human-centric perspective; you’re assuming aliens live like us, and that’s what I meant by the distributions—you’re assuming that the distributions
Adel Daoud: On the conditional distribution given the image space—with M_X denoting the image from Planet X—that conditional distribution doesn’t change much; it’s approximately the same. Thus, if that assumption holds, you can reason accordingly, giving you a reasonable chance of proceeding along these lines. You could employ supervised machine learning in the vein that Connor suggested, which I term “risky business” here, assuming this equality holds fully; then, you can transfer directly.
Moreover, if you assume these distributions are exactly equal—meaning what happens on Earth maps precisely to the other planet—then no training is needed; the only requirement is images from that planet, allowing full transfer of our knowledge to that instance.
You can also utilize unsupervised machine learning models to categorize living conditions on that planet, summarizing the high-dimensional image space into what we call K low-dimensional labels, perhaps indicating richer or poorer extraterrestrial living conditions. This approach aids the task, but critically, as evident here, we need the images—hence, one key takeaway is to prioritize work on the image side, assuming extraterrestrial data will become available soon. However, what we have overarching demonstrated, at least for our planet, is that this is feasible; thus, on Planet Earth, the question—as I alluded to—is to what extent it can be applied elsewhere.
Therefore, while preparing this lecture, I pondered what else this might represent, as I mentioned: the inception of a new unresolved list of problems. There are unresolved problems in extraterrestrial contexts specifically, but none on measurement therein, so I thought we need to add that to the Wikipedia page, at least after this talk.
A big thanks to everyone here, and also to the team on Planet Earth, but obviously we must thank the satellites orbiting Earth that make this possible. Unfortunately, Landsat 6 was killed in action, exploding on launch, but that happens in this business; Landsat 9 joined just two or three years ago in this endeavor. If interested in this work, please talk to me, Ji, Connor, and James; check out our websites. A shout-out for The Journeys of Scholars—it’s a totally different endeavor but also related, featuring interviews with esteemed scholars, and Ji is one of them; Gary King at IQSS is another, discussing the meta-conditions of doing research. Thank you very much.
Audience Member: I do have questions, but before I jump into that, any from anybody here or anybody online? Probably not, so let me ask my question. Oh, there’s a question: a few slides back, where you were showing the predicted actual results, it looked like there were still some trends left over. And then that, plus the map of Africa where you saw by country what the correlation was—do you have a sense of whether it’s some countries that just have much better data than others, and that’s why you have such better predictions for some countries? And does that map back to your predicted-versus-actual graph, where those extra trends that are remaining are because some countries are just systematically not predicted well? What do you think’s going on?
Adel Daoud: It’s really one of the most important questions that we’re thinking about: when you’re looking at that graph, what’s happening in these countries?
Adel Daoud: For Egypt, I already provided an explanation for our poor performance there; another factor involves potential fast-moving processes that we are not capturing. War and conflict remain, unfortunately, integral to Africa’s history, with much of it persisting today, thus some of that dynamic may contribute, though it might not represent the strongest argument for all trends observed.
A further line of thinking, which clearly demands additional qualitative and complementary data for verification, posits that certain countries align their living conditions differently with surface-level changes on Earth, resulting in varied joint distributions. In an ideal scenario, greater construction activity corresponds more closely to alterations in satellite imagery, correlating with economic growth; the stronger this alignment, the more effectively an image serves as a proxy.
However, constructing mere statues, for instance, alters the image without substantially advancing predictive value. As illustrated by nightlight images of North Korea, Pyongyang appears illuminated, yet exhibits minimal temporal change due to statue-building and similar activities, complicating the model’s utility since these nations pursue divergent development paths. Essentially, consider this as a joint distribution where the image proxies the outcome, and the quality of matching between image and outcome determines efficacy—if misaligned, the image underperforms, presuming an appropriate model, which constitutes yet another consideration.
Audience Member: I believe there’s an intriguing connection between sociology and physics, given the strong global correlation between energy consumption and living conditions; therefore, I wonder if employing that energy factor as a bridge between sociological and physical realms could facilitate extension to the extraterrestrial domain.
Adel Daoud: That’s a really interesting thought. To elaborate slightly, how are you envisioning this, and do you have ideas for measurement? I believe the answer is affirmative; I’m pondering how we might quantify it beyond nightlight data, so do you have any suggestions?
Audience Member: Indeed, it would hinge somewhat on the technologies employed, as each energy generation method produces distinct signatures regarding pollution, heat, and related aspects. To a degree, this might be encompassed by varying energy production modes, yet a more general or universal approach to detecting those signatures across diverse technologies could prove beneficial—considering extraterrestrials might utilize ultra-advanced systems unknown to us or remain at a Bronze Age level. If we construct a model spanning societal evolutionary stages, leveraging industrialization as key energy injections.
Adel Daoud: Yes, that reminds me: Avi has been pursuing this approach. You know, he’s the one searching for it; they attempt to detect what one might term space trash or signals from destroyed civilizations. Nevertheless, efforts still primarily focus on fast-moving objects, as that’s the predominant recent strategy, though numerous papers endeavor to identify signatures from potential civilizations in the
Although the past may get destroyed, civilizations could leave behind certain traces. One example that comes to mind relates to the atmospheric question; if there is significant industrialization, consider the carbon dioxide emitted, which alters the atmospheric composition of chemicals and serves as an indicator for the existence of ET, even without high-resolution imagery, though it would not capture intra-planetary variations since such changes would be evenly spread.
Another line of thought involves oil extraction on our planet, which is often accompanied by specific heating signatures corresponding to energy extraction, if not direct usage; although I am not familiar with the oil extraction process, there are towers burning gases, creating a clear signature that would not exist without such activity. Nevertheless, the challenge remains how to measure this on the planet’s surface, which I think is a terrific question and represents the core problem.
One more question is whether there is any way to deal with information regarding something like a political system; perhaps we should return to the North Korea example, or simply state that for DTS or Communist systems, the distribution might be constant in time and space. I have actually thought about that particular example: what if ET were highly egalitarian, sharing resources extensively? In that case, the signature of a city might reflect equality in housing distribution, resulting in uniform heights and densities; in the extreme case of egalitarian ET societies, this would be evident, though it might not directly correspond. Flipping the question, how would a dictatorship appear? In North Korea’s NightLight data,
Pyongyang is essentially the only illuminated city at night; I had included but removed a slide on this, thinking it tangential, but it is actually interesting. Indeed, this could inspire a new sequel to the ET movie: the Communist ET, which somebody should develop. But seriously, returning to Earth, can we use simulations to demonstrate political systems, as this concerns living conditions? Leaders may claim they are not dictators, but physical evidence can contradict that; is it possible to apply this in political science? Why not?
I particularly like the Communist example, as it provides compelling physical evidence of dictatorship or something similar. Yes, exactly; while there is much other evidence, this offers a historical perspective—for instance, can we observe this historically? It is not a bad idea, and I had not thought about it that way until preparing this talk, from which I learned a lot. Let me see if I can show you; do you see this? I can make it disappear. Oh, you just made it completely gone, so you see this.
Adel Daoud: This is a particular satellite image of planet Earth night lights. This is how it looks. To your question about how an egalitarian extraterrestrial society would appear, it is really hard to see exactly from this particular image, but I would say you would see more of an equal distribution of land use, assuming all land is usable. That would be one takeaway just looking at the United States, which is not the most egalitarian society—not yet, exactly. If you look at North and South Korea in 1992 and 2008, you see basically Pyongyang just sitting there, and Seoul shows a huge difference.
It is a large difference in terms of how Seoul is developed compared to Pyongyang, being probably one of the more extreme dictatorships on this planet. Now, for extraterrestrials, I think to some extent we will see similar patterns assuming a similar distribution of how that system would play out; you would see something like this. But interestingly, if you look at most human countries, there is still a concentration of power, with the capital or big city still dominating—literally the concentration of power.
There is a concentration of economic power, cultural power, and political power, and we see this also in Seoul. Thus, we look at North Korea and say they are building statues there but not much else, but in South Korea there is a lot of activity and pride about what Korea has achieved. However, there is a concentration of power towards Seoul, and you can clearly see that.
Audience Member: What year is the world—the 2008?
Adel Daoud: No, that one, I believe, is around 2010 or something like that.
Audience Member: Yeah, were you thinking about something specific there? I was looking at the India map, and when talking about concentration of power, I don’t think that represents what it looks like now.
Adel Daoud: Yes, because the Gangetic plain has more people than almost any country except China—interesting, and right there in this one it looks almost dark. Indeed, this is one image fitted to this particular screen, so you get a lot meshed up. If we do a higher resolution night light image, you will see a better view of what is going on within India. If you look at Africa, which is our target, there is just not a lot of illumination except for South Africa.
But this type of data has been used for about two or three decades to measure income. We use night light in our models, as you remember, as one of the important sources, but we add the daylight component to it to boost the predictions. If you only use this data, you would get an R-squared of around 0.45; adding the daylight, you can boost it up to the numbers we have shown—0.7 and up. Therefore, it gives you some food for thought about how we measure stuff here on this planet and how it potentially could generalize to other planets. It was really fun to compose this talk, at least for me, and I hope that you enjoyed this particular train of thought.
Organizer: Thank you, thank you, thank you. Well, the most important lesson I learned today is that astronomers look up and we look down—that is the key difference. It all depends on which angle you want to look at.