Behavior as a Modality
Doctoral Thesis · IIIT Delhi & University at Buffalo

Behavior As A Modality

A Framework To Enable Automated Persuasion

On explaining, predicting, and optimizing human behavior at the intersection of communication theory, behavioral science, and artificial intelligence.

by Yaman Kumar Singla

What this thesis asks

Can we build a general model of human behavior — one that explains it, predicts it, and optimizes it across platforms and contexts?

Behavioral science has long split into two cultures: those who explain behavior with interpretable mechanisms, and those who predict it from data. This thesis argues that behavior is a modality of communication — and that modeling it jointly with content lets a single family of models do both. Four questions organize the work:

01

Explanation

What strategies and triggers drive human responses to content — and can we make those mechanisms interpretable?

02

Prediction

Can models predict behavior by jointly modeling content and response across the full communication process?

03

Understanding

Can one modality — behavior — improve a model's understanding of another — the content itself?

04

Generation

Can we change content to steer behavior — generating memorable, engaging, persuasive messages by design?

Chapter 1 · Introduction

What is Behavior?

Introducing the two cultures of behavioral sciences

Superscript numbers mark citations in the original text; the full bibliography lives in the thesis PDF.

Behavior as a modality* occurs in the process of communication. Communication includes all of the procedures by which one mind may affect another1. This includes all forms of expression — words, gestures, speech, pictures, and musical sounds. Communication can be seen as being composed of seven modalities: the communicator, the message, the time of the message (or time of receipt), the channel, the receiver, the time of effect, and the effect.

* A modality is a medium through which information is conveyed; a multimodal distribution, similarly, is one with more than one peak. Behavior is treated here as one such medium — a channel of information in its own right, alongside text, image, and audio.

The seven factors of communication: communicator, message, time of message, channel, receiver, time of effect, and effect — illustrated with an Adobe email campaign and its receiver effects (opens, clicks, purchases).
The seven factors of communication. Any message is created to serve an end goal. For marketers, that goal is a desired receiver effect — clicks, purchases, likes, retention. The figure traces the pipeline from communicator to effect.

The seven modalities interact in different ways depending on the context. Four cases illustrate the range:

Mass communication — a speech

The communicator is an individual or organization; the channel is radio, television, or the internet; the effect is the audience's subsequent behavior. The communicator controls three factors — the message, the channel, and the timing. A politician speaking on "Artificial Intelligence" tailors the message to public radio versus a live rally, and a speech near Independence Day carries different themes than one at New Year's. In live delivery, the time of effect is simultaneous with the message.

Social media

Communication becomes asynchronous. The communicator still controls message, channel, and timing, but reach and effect are also shaped by the platform's algorithms. A company announcing a launch on Twitter crafts the message differently than in a live keynote, and posts during peak hours for visibility. A tweet can be shared and rediscovered over time, so the time of effect is no longer simultaneous with delivery.

Peer-to-peer

Two individuals exchange messages over a mutually agreed channel — WhatsApp, email, a phone call. Unlike mass communication, the sender does not choose the audience; the interaction happens only once both parties establish a connection. The effect is more controlled, but still depends on timing, the receiver's interpretation, the sender, and the channel.

Bidirectional exchange

A continuation of the peer-to-peer case where communicator and receiver swap roles each turn. The channel stays the same; the effect of one turn becomes the message of the next, its receiver becoming the next communicator — a conversation.

These modalities vary independently of each other and carry signals about one another2. The message carries information from communicator to receiver; behavior — the effect — carries information back from the receiver. This is often a continuous cycle, where behavior generated in one turn becomes the message of the next, forming a conversation.

Two cultures: explanation and prediction

Different fields of behavioral science deal with different parts of behavior, but two streams have emerged broadly: the explanation and the prediction of behavior3.

Historically, behavioral social scientists have sought explanations that provide interpretable causal mechanisms. Milgram's and Asch's experiments explained obedience to authority4. Cialdini identified six principles of persuasion — reciprocity, commitment, social proof, authority, liking, and scarcity — a framework that explains why certain messaging works5. In economics, Kahneman and Tversky's prospect theory explains decisions under uncertainty through biases like loss aversion6.

This approach of theorizing has worked remarkably well in the physical sciences, where data is plentiful and theories make unambiguous predictions. Newton's laws precisely predict planetary orbits, letting us calculate when Halley's comet returns — every 76 years. Einstein's relativity predicted the bending of light, confirmed during the 1919 solar eclipse. The periodic table predicted undiscovered elements; Maxwell's equations predicted radio waves decades before Hertz demonstrated them.

Cialdini's principles explain why an authority figure influences behavior — but they cannot predict, with Newton-like precision, whether a specific endorsement will lift sales by 15% or 50%.

Such theoretical success has not been replicated in predicting social outcomes7. Human behavior is far more complex and context-dependent than planetary motion. Studies repeatedly show that expert opinions fare no better than non-experts at predicting economic and political trends, societal change, or advertising success — and that non-expert predictions of behavior (which cascades spread, which images are memorable) are roughly as good as a coin toss8. Causal mechanisms still have their merits: they help human decision-makers make intuitive sense of a situation and choose the next action.

In parallel, the availability of behavioral data at scale has drawn machine learning into classically behavioral questions — persuasion strategies, information diffusion, and the predictability of behavior9. This prediction-oriented approach mirrors deep learning's success elsewhere: models classify images with superhuman accuracy (97.8% on ImageNet versus 94.9% for humans) without understanding why an image contains an object; language models achieve remarkable performance through pattern recognition; recommender systems predict preferences — Netflix's was estimated to save the company \$1 billion a year — without modeling the psychology behind them.

Within the prediction community, subfields have splintered: personalization optimizes the receiver for a message; recommendation chooses content for a receiver; effect-prediction forecasts click-through, cascades, sales, and memorability. Each factor of communication is studied in isolation, without relying on the underlying unity of the communication process. Some of the major problems studied across behavioral science:

Sender space

  • Source optimization — who should send a message to a given audience.

Receiver space

  • Personalization
  • Customer segmentation
  • Social network analysis
  • Lookalike modeling
  • Market surveys
  • Identity stitching
  • Behavior explanation

Content space

  • Recommender systems
  • A/B testing
  • Customer targeting
  • Propensity / engagement modeling
  • Transsuasion & transcreation
  • Search engine optimization
  • Performant content generation
  • Argument mining · persuasion strategies

Channel space

  • Channel optimization
  • Marketing-mix modeling
  • Auction design & bidding

Time space

  • Send-time optimization
  • Trend forecasting

A common theme runs through all of it: the intent to control behavior — lift click-through from 2% to 5%, boost turnout by 15%, double engagement. From that intent, explanation and prediction serve as intermediate steps toward control and optimization. Optimizing behavior means fulfilling the communicator's objectives by strategically managing the other six parts of the communication process — the right spokesperson, message, channel, timing, and audience — and measuring the outcome.

But current approaches are fragmented: a model trained to predict Twitter engagement cannot predict YouTube views; an ad optimizer for fashion fails on food. The solution requires a general understanding of behavior that transfers across domains, platforms, and contexts.

The digital age, and a lesson from language models

The digital age is marked by human behavioral data in huge repositories — data that is big, always-on, and observational, but also incomplete and algorithmically confounded10. Prior predictive work leaned on individual platforms — Twitter, Instagram, Google Trends, Wikipedia, shopping sites — but stayed limited to one platform, one question, one user type. We want a model that understands human behavior in general, not one effect on one platform for one kind of user.

This parallels natural language processing, where supervised models were limited by available supervision and could answer only the one question they were trained for. The field solved it with Large Language Models — general-purpose models that understand language and can do sentiment analysis, question answering, translation, and more, zero-shot. Two things have always worked for neural networks: larger models and more data. Going from millions of tokens to trillions increased transfer across a wide variety of tasks.

Levels of content analysis arranged in a hierarchy, from surface features up to the receiver's behavioral effect. Humans predict the first three levels well but not the last.
Levels of content analysis. Tasks arranged in a hierarchy roughly based on the levels of language. Notably, humans are good at the first three levels — but not the last, the level of behavioral effect.

So how do we build a model that understands behavior in general? We take the LLM recipe and apply it to the behavioral repositories on the internet, whose format is exactly the seven-factor communication model. Because those repositories are incomplete, not every factor is always present — but a subset always is, and scale plus a large model yields a general behavior-understanding model. We call it the Large Content and Behavior Model (LCBM), and show it can predict behavior, explain it, and generate messages to bring about behavior11.

Are general LLMs already able to solve behavioral problems? We test GPT-3.5, GPT-4, and Llama models — and find they cannot. The reason is structural: LLMs model only one factor (the message) out of the seven, treating the communicator, receiver, channel, time, and behavior as "noise" to be purged from training data. That systematic removal is why the models develop no behavioral capabilities. Even multimodal models like LLaVA, after training on hundreds of thousands of instructions, can "see" — but only answer questions at the first two levels of content analysis, because their alignment data lives there too.

Diagram linking the thesis chapters across the two pillars of understanding/explanation and prediction, over the seven factors of communication.
How the chapters connect. Following the two traditions of behavioral science, the thesis moves through explanation and prediction — and how each chapter links to the next across the seven factors of communication.
This is the argument in brief. The full Introduction — with every example, citation, and figure — is in the thesis.
Read the PDF

The body of the thesis

Four chapters, from explaining behavior to generating it

The Introduction and Conclusion are on this page. The four chapters between them — the technical core — are summarized here and covered in full in the PDF.

  1. 2

    Explaining Behavior: Persuasion Strategies & Universal Adversarial Triggers

    The largest taxonomy of persuasion strategies in image and video ads, plus universal adversarial triggers that probe what behavioral models actually learn.

  2. 3

    Modeling Behavior: A Case for Large Content and Behavior Models

    LCBMs model content and the behavior it induces. Trained on 40,000+ YouTube videos and 168M tweets, they show emergent few-shot and zero-shot behavior.

  3. 4

    Analyzing Behavior: Teaching Behavior Improves Content Understanding

    Encoding behavioral signals improves content understanding across 46 tasks and 23 datasets in language, audio, text, and video — with no architectural changes.

  4. 5

    Optimizing Behavior: Generating Content to Optimize Behavior

    Henry generates more memorable ads (+44%); EngageNet and Engagement Arena generate and benchmark images for real engagement, not just aesthetics.

Chapter 6 · Conclusion

Conclusion & an Outlook for Future Work

Where the framework arrives, and the questions it opens

This thesis explored the intersection of communication theory, behavioral science, and artificial intelligence — explaining, understanding, and optimizing human behavior through large-scale modeling. The work builds on the seven-factor model of communication while leveraging unprecedented access to digital behavioral data to advance both explanatory and predictive approaches.

In persuasion-strategy analysis, we developed the most extensive framework of generic persuasion strategies to date, and released the first datasets for studying them in image and video advertisements. We showed that existing LLMs are inherently limited in modeling behavior — because behavioral data is systematically removed during training — and built LCBMs that integrate all seven factors of communication, releasing behavior instruction-tuning data from 40,000+ YouTube videos and 168M tweets. We demonstrated that behavioral signals enhance content understanding across 46 tasks and 23 datasets, and in generation we built Henry (a 44% improvement in memorability) and EngageNet with Engagement Arena, the first automated benchmark for the engagement potential of text-to-image models.

An outlook: automated behavioral science

The integration of behavioral data into AI opens several directions for more nuanced, context-aware models of human response:

Infinite personalization

The printing press solved production; steam and the internet solved delivery. The last limiting factor is the human cost of producing performant content — which is why we need models that generate it, enabling a personalized channel between any communicator and receiver.

Digital humans & digital societies

LLMs conditioned on demographic and psychological profiles can act as "silicon samples" of populations — digital twins that serve as testing grounds for interventions, policy, and communication strategy, if we ensure they authentically represent the diversity of human experience.

Measuring AI persuasion

As models learn to generate verifiably persuasive content, we must rigorously study, measure, benchmark, and monitor their persuasive capabilities — for advertising and social good, and against misinformation and manipulation.

Automatically explaining behavior

Prediction and explanation are drifting apart, yet human curiosity wants the mechanism. The frontier is methods that carry both high predictive power and scalability — bridging the two cultures with simulation and natural experiments.

Rethinking free will

One practical definition: free will is the lack of predictability in human actions. If a model can predict your next action with high accuracy, the existence of true free will comes into question.

If a model is able to predict an individual's next action to a high degree of accuracy, then the existence of true free will comes into question.

Societal & ethical implications

Behavior-optimized AI is fundamentally dual-use: the technologies that benefit education, public health, and design can also be misused for political manipulation, misinformation, and consumer exploitation. Across the thesis we implemented safeguards — acceptable-use policies, PII removal, aggregation, and controlled evaluation — but the hard problems remain open: how to quantitatively measure a model's manipulation potential; how to protect privacy at behavioral scale; how to ensure fairness across cultures and demographics; and what informed consent even means when a system can predict and influence future actions. Responsible development demands transparency, participatory design, continuous monitoring, and adaptive governance — the power to understand and influence behavior must be wielded in service of human flourishing.

Open research questions

The thesis closes on the questions it cannot yet answer, grouped into four thematic areas:

Modeling & architecture

  • Fusing text, image, and behavior in one model
  • Long-term, sequential behavioral dynamics
  • Data-efficient modeling in low-data domains

Causality & intervention

  • Integrating causal inference with prediction
  • From passive prediction to automated intervention
  • Persuasive, trustworthy explanations

Ethics & society

  • "Constitutional AI" for persuasion
  • Quantifying & mitigating manipulation risk
  • Building & validating digital societies

Theory, practice & people

  • Cross-cultural & demographic generalization
  • Theory-driven, interdisciplinary modeling
  • Real-world validation & deployment

Finally, two thoughts worth remembering as we enter what may be the fourth major phase in the study of communication:

We're actually much better at planning the flight path of an interplanetary rocket (rocket science) than we are at managing the economy, merging two corporations, or even predicting how many copies of a book will sell (behavior prediction). So why is it that rocket science seems hard, whereas problems having to do with people — which arguably are much harder — seem like they ought to be just a matter of common sense?

— Duncan J. Watts

Nothing in Nature is random. A thing appears random only through the incompleteness of our knowledge.

— Baruch Spinoza

A closing reflection

As we confront the scientific challenges to the notion of free will, it is worth recognizing that these questions have been explored in philosophical and religious traditions for millennia. The Bhagavad Gita offers a profound perspective on agency — suggesting that what we perceive as free will may be shaped by forces beyond our conscious control. This resonates with modern behavioral science and neuroscience, which increasingly point to the limits of conscious agency. In Chapter 3, Verse 27, the Gita states:

Prakṛteḥ kriyamāṇāni guṇaiḥ karmāṇi sarvaśaḥ;
Ahaṅkāra-vimūḍhātmā kartāham iti manyate.

"All actions are performed by the modes of material nature, but a person deluded by false identification with the ego thinks, 'I am the doer.'"

These verses articulate a philosophy of action that acknowledges the limits of personal agency and encourages detachment from the fruits of action. The Gita's perspective is not one of fatalism, but an invitation to act with awareness of the larger forces at play — recognizing that while we must act, we are not the sole authors of our actions. This ancient wisdom aligns with contemporary scientific insight into the nature, causes, and limits of human agency.

Read the whole thing

The complete thesis — six chapters, the datasets, the models, the proofs.

You've read the bookends. The full argument, every experiment, and the complete bibliography are in the PDF.