Digital platforms record an enormous amount of human activity. People search for information, respond to posts, follow news events and leave behind behavioral traces that can be counted at a scale conventional surveys rarely reach.
The difficulty is interpretation. A surge in searches does not necessarily represent public opinion. A heavily discussed position may be driven by a small number of unusually active users. News coverage can show what editors considered important without necessarily showing what audiences were trying to understand.
These measurement problems have become a recurring theme in the work of Mohamed Soufan, a software engineer and computational researcher who has been using digital data to study public attention and behavior, particularly in Lebanon.
Rather than treating online activity as a direct representation of society, his research asks a narrower question: what can a particular digital signal actually tell us, and what becomes visible when it is compared with another one?
From Software Engineering to Social Measurement
Computational social research often sits between two very different kinds of work. One involves questions about behavior, attention and social conditions. The other involves the practical problem of turning messy information into something that can be measured.
Soufan approaches these questions from a software engineering background. That matters less as a biographical detail than as a methodological one. Many of the problems involved in studying digital behavior resemble engineering problems: defining inputs, deciding how observations should be represented, identifying failure cases and checking whether an output actually corresponds to what the system claims to measure.
A metric such as replies, for example, may look straightforward until it is used as evidence of engagement. Replies, likes and reposts are all interactions, but they involve different actions and motivations. Search activity presents a similar problem. It can indicate demand for information without revealing exactly why that demand exists.
Framing such questions as measurement problems makes the assumptions behind familiar metrics easier to examine.
Building Data Where No Dataset Exists
The analytical work often starts before there is a usable dataset.
News articles, search records, public reports and platform activity may all contain relevant information, but they are usually produced for purposes other than research. They have to be collected, filtered and converted into comparable observations before meaningful analysis can begin.
That process is not neutral. Decisions about which records count, how categories are defined, when two reports describe the same event or how ambiguous cases are handled can affect the eventual result.
In this kind of research, data construction therefore becomes part of the methodology rather than a technical step that can be separated from it.
Soufan’s projects have involved building custom pipelines around specific research questions, structuring raw records and checking intermediate outputs before using them in comparisons. The approach is useful when the information needed to answer a question exists publicly but has never been assembled in a form designed for analysis.
It also creates a constraint: conclusions are only as reliable as the definitions and transformations used to produce the underlying dataset.
When Different Signals Tell Different Stories
Comparing datasets becomes particularly useful when they do not agree.
News attention, search behavior and platform participation may all respond to the same event, but they measure different forms of activity. Treating one as a substitute for another can obscure those differences.
In Soufan’s research on Lebanon, international media attention and local search behavior did not always emphasize the same issues. Other work examined how visible online discussion can be disproportionately shaped by highly active participants, while research into platform engagement considered why replies should not automatically be treated as equivalent to likes or reposts.
The significance of these comparisons is not that one dataset must be correct and another wrong.
A news dataset can accurately measure coverage while a search dataset simultaneously measures information demand. Both may describe the same period accurately while producing very different pictures of it.
That difference can itself become analytically useful. It can distinguish visibility from prevalence, expression from information-seeking and attention from demand.
Instead of asking which source represents society, the more precise question is what each source represents.
Lebanon as a Difficult Test Case
Lebanon has featured repeatedly in this work partly because it is a difficult environment in which to measure social change.
Economic conditions, migration concerns, security events and public-service pressures can shift quickly, while relevant information may be dispersed across news organizations, institutions and digital platforms. Conventional statistics can also arrive after the behavior researchers are interested in has already changed.
Digital data offers more immediate signals, but immediacy creates its own problems. Major events can dominate media coverage. Online participation may be concentrated among particular groups. Search activity reflects people with internet access and cannot identify the intentions of every person represented in an aggregate trend.
These conditions make Lebanon less a source of clean answers than a useful test of whether a measurement remains informative under pressure.
One of Soufan’s studies compared international news coverage of Lebanon during the March 2026 conflict with Lebanon-related search interest. The two distributions differed substantially, with conflict accounting for a much larger share of the classified news coverage than of search interest. Searches continued to reflect concerns related to subjects including living conditions, the economy and migration.
The comparison did not establish that one side represented the public better than the other. It showed that media attention and information demand were measuring different things during the same period.
From Measuring Attention to Detecting Pressure
A broader question follows from these comparisons: can digital signals be useful before a change becomes clearly visible in conventional indicators?
A single spike in migration-related searches would provide weak evidence on its own. The same is true of an isolated movement in prices, mobility or demand for a public service.
The possibility becomes more interesting when independent sources begin changing at the same time.
In a recent interview with Inside Telecom, Soufan discussed the potential value of connecting data that is often held or analyzed separately. Search behavior, mobility patterns, pricing information, public-service records and media activity can each capture a different aspect of changing conditions.
The idea is closer to early detection than prediction.
Multiple signals moving together might indicate that something deserves investigation before it appears clearly in quarterly statistics or formal assessments. For businesses, that could mean noticing a shift in consumer pressure or demand. For public institutions, it could mean identifying growing strain on a service. Humanitarian organizations may be interested in changing mobility or information-seeking patterns.
None of those signals would be sufficient on its own. Their usefulness would depend on whether independent sources corroborate rather than merely repeat one another.
What Digital Data Cannot Establish
The availability of large datasets can make conclusions appear more certain than they are.
Search trends can show changes in relative interest but cannot explain every user’s motive. Social-media activity can measure what participating users said or did without establishing that they represent the broader population. News databases can quantify coverage without determining how socially important an issue was.
Scale does not resolve those problems.
A dataset containing millions of observations may still provide a biased view if the underlying population or behavior is systematically different from what the researcher is trying to understand. Associations can identify patterns without establishing causation, while sudden changes may warrant investigation without explaining what produced them.
These limitations are particularly important when computational methods are applied to social questions because the variables often stand in for concepts that cannot be observed directly.
The defensible use of digital traces is therefore narrower than the claim that online data can somehow provide a real-time picture of society. Their value lies in providing additional measurements that can be compared with other evidence.
Asking Better Questions of the Data
The amount of behavioral data available to researchers is unlikely to be the main constraint in the coming years. The harder problem will remain deciding what those records actually measure.
Searches, headlines, replies, mobility records and other digital traces capture different actions generated for different reasons. Combining them does not automatically produce a more accurate picture. The analytical value comes from defining the relationship between the signal and the question being asked.
That is the common thread running through Soufan’s recent research. The objective is not to find a single dataset that represents society, but to make narrower measurements, compare them carefully and pay attention when apparently related indicators diverge.
In computational research, more data can expand what is observable. It does not remove the need to decide what the observations mean.



