7 References
[1]
Aher, G.V. et al. 2023. Using large language
models to simulate multiple humans and replicate human subject studies.
Proceedings of the International Conference on Machine
Learning. (2023).
[2]
Aiello, L.C. and Dunbar, R.I.M. 1993. Neocortex
size, group size, and the evolution of language. Current
Anthropology. 34, 2 (1993), 184–193.
[3]
Argyle, L.P. et al. 2023. Out of one, many:
Using language models to simulate human samples. Political
Analysis. 31, 3 (2023), 337–351.
[4]
Axelrod, R. 1984. The evolution of
cooperation. Basic Books.
[5]
Bai, C. et al. 2021. M2p2: Multimodal
persuasion prediction using adaptive fusion. IEEE Transactions on
Multimedia. (2021).
[6]
Bai, H. et al.
2023. The potential of generative AI for personalized
persuasion at scale. Scientific Reports. (2023).
[7]
Baldwin, J.M. 1896. A new factor in evolution.
The American Naturalist. 30, 354 (1896), 441–451.
https://doi.org/10.1086/276408.
[8]
Banerjee, S. and Lavie, A. 2005. METEOR: An
automatic metric for MT evaluation with improved correlation with human
judgments. Proceedings of the acl workshop on intrinsic and
extrinsic evaluation measures for machine translation and/or
summarization (2005), 65–72.
[9]
Bankes, S.C. 2002. Agent-based modeling: A
revolution? Proceedings of the National Academy of Sciences.
99, suppl_3 (2002), 7199–7200.
[10]
Barrett, M. et al. 2016. Cross-lingual transfer of
correlations between parts of speech and gaze features.
Proceedings of COLING 2016, the 26th international
conference on computational linguistics: Technical papers (Osaka,
Japan, Dec. 2016), 1330–1339.
[11]
Bates, E. and MacWhinney, B. 1989.
Functionalism and the competition model. The crosslinguistic study
of sentence processing. B. MacWhinney and E. Bates, eds. Cambridge
University Press. 3–73.
[12]
Baumgartner, J. et al. 2020. The pushshift
reddit dataset. Proceedings of the international AAAI conference on
web and social media (2020), 830–839.
[13]
Belen, R.A.J. de et al. 2022. ScanpathNet: A
recurrent mixture density network for scanpath prediction. IEEE
CVPR (2022).
[14]
Berzak, Y. et al. 2022. CELER: A
365-participant corpus of eye movements in L1 and L2 english reading.
Open Mind. (2022).
[15]
Bhattacharyya, A. et al. 2023. A video is worth
4096 tokens: Verbalize videos to understand them in zero shot.
Proceedings of the 2023 conference on empirical methods in natural
language processing (2023), 9822–9839.
[16]
Bhattacharyya, A. et al. 2026.
ALPHA: Aligning LLMs with ad engagement data.
Proceedings of the AAAI conference on artificial intelligence
(2026).
[18]
Bhattacharyya, S. et al. 2025. Unsupervised
large-scale memorability modeling from tip-of-the-tongue retrieval
queries. arXiv preprint arXiv:2511.20854. (2025).
[19]
Bickerton, D. 1983. Creole languages.
Scientific American. 249, 1 (1983), 108–115.
[20]
Bickerton, D. 1990. Language and
species. University of Chicago Press.
[21]
Boyd, R. and Richerson, P.J. 1985. Culture
and the evolutionary process. University of Chicago Press.
[22]
Breakthrough 2023. PySceneDetect:
Video scene cut detection and analysis tool. GitHub.
[23]
Broockman, D. et al. 2024. Can
LLMs replace human evaluators? Predicting the results of
social science experiments. arXiv preprint arXiv:2406.14508.
(2024).
[24]
Brown, T.B. et al.
2020. Language models are few-shot learners. Advances in Neural
Information Processing Systems. 33, (2020).
[25]
Byrne, R.W. and Whiten, A. 1988. Machiavellian
intelligence. Behavioral and Brain Sciences. 11, 2 (1988),
233–244.
[26]
Cheney, D.L. et al. 1995. The role of grunts in
reconciling opponents and facilitating interactions among adult female
baboons. Animal Behaviour. 50, (1995), 249–257.
[27]
Cheney, D.L. and Seyfarth, R.M. 1997.
Reconciliatory grunts by dominant female baboons influence victims’
behaviour. Animal Behaviour. 54, (1997), 409–418.
[28]
Chiang, W.-L. et
al. 2024. Chatbot arena: An open platform for evaluating llms by
human preference. arXiv preprint arXiv:2403.04132.
(2024).
[29]
Chiang, W.-L. et al. 2023. Vicuna: An open-source
chatbot impressing GPT-4 with 90%* ChatGPT quality.
[30]
Chung, H.W. et al. 2022. Scaling instruction-finetuned
language models.
[31]
Clutton-Brock, T.H. et al. 1999. Selfish
sentinels in cooperative mammals. Science. 284, 5420 (1999),
1640–1644.
[32]
Cohendet, R. et al. 2019. VideoMem:
Constructing, analyzing, predicting short-term and long-term video
memorability. Proceedings of the IEEE/CVF international conference
on computer vision (2019), 2531–2540.
[33]
Curtiss, S. 1977. Genie: A psycholinguistic
study of a modern-day wild child. Academic Press.
[34]
Damasio, H. et al. 1996. A neural basis for
lexical retrieval. Nature. 380, 6574 (1996), 499–505.
https://doi.org/10.1038/380499a0.
[35]
Danescu-Niculescu-Mizil, C. et al. 2012. Echoes
of power: Language effects and power differences in social interaction.
Proceedings of the 21st international conference on world wide
web (2012), 699–708.
[36]
Danescu-Niculescu-Mizil, C. et al. 2012. You
had me at hello: How phrasing affects memorability. arXiv preprint
arXiv:1203.6360. (2012).
[37]
Deng, J. et al. 2009. ImageNet: A
large-scale hierarchical image database. Proceedings of the IEEE
conference on computer vision and pattern recognition (2009),
248–255.
[38]
Dennett, D.C. 1995. Darwin’s dangerous
idea: Evolution and the meanings of life. Simon &
Schuster.
[39]
Devlin, J. et al. 2019. BERT:
Pre-training of deep bidirectional transformers for language
understanding. Proceedings of NAACL-HLT. (2019),
4171–4186.
[40]
Du, Y. et al. 2020. PP-OCR: A practical ultra
lightweight OCR system.
[41]
Dunbar, R. 1998. Grooming, gossip, and the
evolution of language. Harvard University Press.
[42]
Dunbar, R.I.M. 1992. Neocortex size as a
constraint on group size in primates. Journal of Human
Evolution. 22, 6 (1992), 469–493.
[43]
Durmus, E. and Cardie, C. 2018. Exploring the role of prior
beliefs for argument persuasion. NAACL: Human language
technologies, volume 1 (long papers) (New Orleans, Louisiana, June
2018), 1035–1045.
[44]
Falk, E.B. et al. 2010. The neural correlates
of persuasion: A common network across cultures and media.
Journal of Cognitive Neuroscience. 22, 11 (2010), 2447–2459.
https://doi.org/10.1162/jocn.2009.21363.
[45]
Fedorenko, E. et al. 2024. Language is
primarily a tool for communication rather than thought. Nature.
630, (2024). https://doi.org/10.1038/s41586-024-07522-w.
[46]
Ferretti, F. and Adornetti, I. 2021. Persuasive
conversation as a new form of communication in Homo
sapiens. Philosophical Transactions of the Royal Society B:
Biological Sciences. 376, (2021), 20200196. https://doi.org/10.1098/rstb.2020.0196.
[47]
Festinger, L. 1957. A theory of cognitive
dissonance. Stanford University Press.
[49]
Frisch, K. von 1967. The dance language and
orientation of bees. Harvard University Press.
[50]
Gardner, B.T. and Gardner, A.R. 1974. Comparing
the early utterances of child and chimpanzee. Minnesota symposium on
child psychology. A. Pick, ed. University of Minnesota Press.
3–23.
[51]
Gellner, E. 1988. Plough, sword and book:
The structure of human history. University of Chicago Press.
[52]
Gerber, A.S. et al. 2016. A field experiment
shows that subtle linguistic cues might not affect voter behavior.
Proceedings of the National Academy of Sciences. 113, 26
(2016), 7112–7117.
[53]
Gopnik, M. 1990. Feature-blind grammar and
dysphasia. Nature. 344, (1990), 715.
[54]
Gopnik, M. and Crago, M.B. 1991. Familial
aggregation of a developmental language disorder. Cognition.
39, 1 (1991), 1–50. https://doi.org/10.1016/0010-0277(91)90058-C.
[55]
Habernal, I. and Gurevych, I. 2016. What makes
a convincing argument? Empirical analysis and detecting attributes of
convincingness in web argumentation. Proceedings of the 2016
conference on empirical methods in natural language processing
(2016).
[56]
Hackenburg, K. and Margetts, H. 2024.
Evaluating the persuasive influence of political microtargeting with
large language models. Proceedings of the National Academy of
Sciences. 121, 24 (2024).
[57]
Hadoux, E. et al. 2021. Strategic argumentation
dialogues for persuasion: Framework and experiments based on modelling
the beliefs and concerns of the persuadee. arXiv preprint
arXiv:2101.11870. (2021).
[58]
Hamilton, W.D. 1964. The genetical evolution of
social behaviour, I and II. Journal of
Theoretical Biology. 7, 1 (1964), 1–52.
[59]
Harari, Y.N. 2015. Sapiens: A brief history
of humankind. Harper.
[60]
Hauser, M.D. et al. 2002. The faculty of
language: What is it, who has it, and how did it evolve?
Science. 298, 5598 (2002), 1569–1579. https://doi.org/10.1126/science.298.5598.1569.
[61]
He, S. et al. 2019. Human attention in image
captioning: Dataset and analysis. ICCV (2019).
[62]
Henrich, J. et al. 2010. The weirdest people in
the world? Behavioral and Brain Sciences. 33, 2–3 (2010),
61–83. https://doi.org/10.1017/S0140525X0999152X.
[63]
Hockett, C.F. 1960. The origin of speech.
Scientific American. 203, 3 (1960), 88–96.
[64]
Hollenstein, N. et al. 2019. Advancing NLP with
cognitive language processing signals. arXiv preprint
arXiv:1904.02682. (2019).
[65]
Hollenstein, N. et al. 2021. CMCL 2021
shared task on eye-tracking prediction. Proceedings of the
workshop on cognitive modeling and computational linguistics
(Online, June 2021), 72–78.
[66]
Hollenstein, N. et al. 2022. CMCL
2022 shared task on multilingual and crosslingual prediction of human
reading behavior. CMCL shared task on multilingual and
crosslingual prediction of human reading behavior (2022).
[67]
Hussain, Z. et al. 2017. Automatic understanding of image
and video advertisements.
[68]
I,
H.S. et al. 2024. Long-term
ad memorability: Understanding and generating memorable ads.
[70]
Jha, P. et al.
2025. Improving image generation for advertising via self-play reward
optimization. Advances in neural information processing systems
(2025).
[71]
Kapoor, S. et al. 2024. Measuring and improving
the well-being of users in recommendation systems. Proceedings of
the ACM on Human-Computer Interaction. (2024).
[72]
Karessli, N. et al. 2017. Gaze embeddings for
zero-shot image classification. IEEE CVPR (2017).
[73]
Khandelwal, A. et al. 2024. Large content and
behavior models to understand, simulate, and optimize content and
behavior. The twelfth international conference on learning
representations (2024).
[74]
Khurana, T. et al.
2023. Behavior optimized image generation via online exploration.
arXiv preprint arXiv:2311.10995. (2023).
[75]
Kirillov, A. et al. 2023. Segment anything.
[76]
Klimt, B. and Yang, Y. 2004. The enron corpus:
A new dataset for email classification research. European conference
on machine learning (2004), 217–226.
[77]
Kramer, A.D.I. et al. 2014. Experimental
evidence of massive-scale emotional contagion through social networks.
Proceedings of the National Academy of Sciences. 111, 24
(2014), 8788–8790. https://doi.org/10.1073/pnas.1320040111.
[78]
Krebs, J.R. and Dawkins, R. 1984. Animal
signals: Mind-reading and manipulation. (1984).
[79]
Krizhevsky, A. et al. 2012.
ImageNet classification with deep convolutional neural
networks. Advances in neural information processing systems
(2012).
[80]
Kumar, M. et al. 2023. Persuasion strategies in
advertisements. arXiv preprint arXiv:2208.09626. (2023).
[81]
Kumar, Y.K. et al.
2024. Large content and behavior models to understand, simulate, and
optimize content and behavior. arXiv preprint arXiv:2410.02653.
(2024).
[82]
Lai, C.S.L. et al. 2001. A forkhead-domain gene
is mutated in a severe speech and language disorder. Nature.
413, 6855 (2001), 519–523. https://doi.org/10.1038/35097076.
[83]
Lasswell, H.D. 1948. The structure and function
of communication in society. The communication of ideas. 37, 1
(1948), 136–139.
[84]
Lazer, D. et al. 2021. Meaningful measures of
human society in the twenty-first century. Nature. 595, (2021),
189–196.
[85]
Lazer, D. et al. 2014. The parable of
Google Flu: Traps in big data analysis. Science.
343, 6176 (2014), 1203–1205.
[86]
Lemmerich, F. et al. 2019. World versus
Wikipedia: Measuring the coverage of the world in language
editions of Wikipedia. Proceedings of the ACM Web
Science Conference. (2019).
[87]
[88]
Li, K. et al. 2023. VideoChat: Chat-centric video
understanding.
[89]
Longpre, L. et al. 2019. Persuasion of the
undecided: Language vs. The listener. Proceedings of the 6th
workshop on argument mining (2019).
[90]
Luke, S.G. and Christianson, K. 2018. The provo
corpus: A large eye-tracking corpus with predictability norms.
Behavior research methods. (2018).
[91]
Machajdik, J. and Hanbury, A. 2010. Affective
image classification using features inspired by psychology and art
theory. Proceedings of the 18th ACM international conference on
multimedia (2010), 83–92.
[92]
Martin, D. et al. 2022. ScanGAN360: A
generative model of realistic scanpaths for 360° images. IEEE
Transactions on Visualization and Computer Graphics. (2022).
https://doi.org/10.1109/TVCG.2022.3150502.
[93]
Maynard Smith, J. and Szathmáry, E. 1995.
The major transitions in evolution. Oxford University
Press.
[94]
McMahan, H.B. et
al. 2013. Ad click prediction: A view from the trenches.
Proceedings of the 19th ACM SIGKDD international conference on
knowledge discovery and data mining (2013), 1222–1230.
[95]
Mikels, J.A. et al. 2005. Emotional category
data on images from the international affective picture system.
Behavior research methods. 37, (2005), 626–630.
[96]
Nowak, M.A. 2006. Five rules for the evolution
of cooperation. Science. 314, 5805 (2006), 1560–1563.
[97]
Nowak, M.A. and Sigmund, K. 2005. Evolution of
indirect reciprocity. Nature. 437, (2005), 1291–1298.
[98]
Ojha, U. et al. 2026. MEMENTO: Web
as a learning signal for low-data advertising domains. arXiv
preprint arXiv:2605.29795. (2026).
[100]
OpenAI 2023. GPT-4 technical
report.
[101]
Papineni, K. et al. 2002. Bleu: A method for
automatic evaluation of machine translation. Proceedings of the 40th
annual meeting of the ACL (2002), 311–318.
[102]
Park, J.S. et al. 2023. Generative agents:
Interactive simulacra of human behavior. Proceedings of UIST.
(2023).
[103]
Pellegrino, G. di et al. 1992. Understanding
motor events: A neurophysiological study. Experimental Brain
Research. 91, 1 (1992), 176–180. https://doi.org/10.1007/BF00230027.
[104]
Pellert, M. et al. 2023. AI psychometrics:
Assessing the psychological profiles of large language models through
psychometric inventories. Perspectives on Psychological
Science. (2023).
[105]
Petrovic, S. et al. 2011. Rt to win! Predicting
message propagation in twitter. Proceedings of the international
AAAI conference on web and social media (2011), 586–589.
[106]
Petty, R.E. and Cacioppo, J.T. 1986. The
elaboration likelihood model of persuasion. Communication and
persuasion. Springer. 1–24.
[107]
Plato 2000. The republic. Cambridge
University Press.
[108]
Pressman, J.D. et al. 2023. Simulacra aesthetic
captions. Technical Report Version 1.0, Stability AI, 2022. url
https://github. com/JD ….
[109]
Qin, X. et al. 2020. U2-net: Going deeper with
nested u-structure for salient object detection. Pattern
recognition. 106, (2020), 107404.
[110]
Raffel, C. et al. 2020. Exploring the limits of
transfer learning with a unified text-to-text transformer. Journal
of Machine Learning Research. 21, 140 (2020), 1–67.
[111]
Ratnieks, F.L.W. 1988. Reproductive harmony via
mutual policing by workers in eusocial Hymenoptera.
American Naturalist. 132, (1988), 217–236.
[113]
Salemi, A. et al. 2023. LaMP: When
large language models meet personalization. arXiv preprint
arXiv:2304.11406. (2023).
[114]
Seeley, T.D. 2010. Honeybee democracy.
Princeton University Press.
[115]
Seeley, T.D. 1995. The wisdom of the hive:
The social physiology of honey bee colonies. Harvard University
Press.
[116]
Seyfarth, R.M. et al. 1980. Monkey responses to
three different alarm calls: Evidence for predator classification and
semantic communication. Science. 210, (1980), 801–803.
[117]
Si, H. et al. 2024. Long-term ad memorability:
Understanding and generating memorable ads. arXiv preprint
arXiv:2309.00378. (2024).
[118]
Singh, A. et al. 2025. FSPO:
Few-shot preference optimization of synthetic preference data in
LLMs elicits effective personalization to real users.
arXiv preprint arXiv:2502.19312. (2025).
[119]
Singh, S. et al. 2025. Teaching human behavior
improves content understanding abilities of LLMs.
International conference on learning representations
(2025).
[120]
Singh, S.K. et al. 2025. Transsuasion:
Introducing and measuring behavior transfer capabilities.
International conference on learning representations
(2025).
[121]
Singla, Y.K. et al. 2022. What do audio
transformers hear? Probing their representations for language delivery
& structure. 2022 IEEE international conference on data mining
workshops (ICDMW) (2022), 910–925.
[122]
Song, C. et al. 2010. Limits of predictability
in human mobility. Science. 327, 5968 (2010), 1018–1021.
[123]
Stab, C. and Gurevych, I. 2014. Annotating
argument components and relations in persuasive essays. Proceedings
of COLING 2014, the 25th international conference on computational
linguistics: Technical papers (2014), 1501–1510.
[124]
Stab, C. and Gurevych, I. 2017. Parsing
argumentation structures in persuasive essays. Computational
Linguistics. 43, 3 (2017), 619–659.
[125]
Stander, P.E. 1992. Cooperative hunting in
lions: The role of the individual. Behavioral Ecology and
Sociobiology. 29, 6 (1992), 445–454.
[126]
Számadó, S. 2010. Pre-hunt communication
provides context for the evolution of early human language.
Biological Theory. 5, 4 (2010).
[127]
Tan, C. et al. 2016. Winning arguments:
Interaction dynamics and persuasion strategies in good-faith online
discussions. Proceedings of the 25th international conference on
world wide web (2016), 613–624.
[128]
Tooby, J. and Cosmides, L. 1990. Toward an
adaptationist psycholinguistics. Behavioral and Brain Sciences.
(1990), 760–762.
[129]
Touvron, H. et al.
2023. LLaMA: Open and efficient foundation language models.
arXiv preprint arXiv:2302.13971. (2023).
[130]
Trivers, R.L. and Hare, H. 1976. Haplodiploidy
and the evolution of the social insects. Science. 191, 4224
(1976), 249–263.
[131]
Voelkel, J.G. et al. 2024. Artificial
intelligence can persuade humans on political topics. PNAS
Nexus. 3, 2 (2024).
[132]
Waal, F. de 1982. Chimpanzee politics:
Power and sex among apes. Jonathan Cape.
[133]
Wachsmuth, H. et al. 2017. Computational
argumentation quality assessment in natural language. Proceedings of
the 15th conference of the european chapter of the association for
computational linguistics: Volume 1, long papers (2017),
176–187.
[134]
Wang, K. et al. 2018. Retweet wars: Tweet
popularity prediction via dynamic multimodal regression. 2018 IEEE
winter conference on applications of computer vision (WACV) (2018),
1842–1851.
[135]
Wei, P. et al. 2022. CREATER:
CTR-driven advertising text generation with controlled
pre-training and contrastive fine-tuning. arXiv preprint
arXiv:2205.08943. (2022).
[136]
Wilber, M.J. et al. 2017. Bam! The behance
artistic media dataset for recognition beyond photography.
Proceedings of the IEEE international conference on computer
vision (2017), 1202–1211.
[137]
Wilson, E.O. 1971. The insect
societies. Harvard University Press.
[138]
Wu, X. et al. 2023. eMotions: A large-scale dataset
for emotion recognition in short videos.
[139]
Xu, H. et al. 2022. GMFlow: Learning optical flow
via global matching.
[140]
Yang, J. et al. 2023. EmoSet: A large-scale
visual emotion dataset with rich attributes. Proceedings of the
IEEE/CVF international conference on computer vision (2023),
20383–20394.
[141]
Yang, J. et al. 2023. Track anything: Segment anything
meets videos.
[142]
Yang, Z. et al.
2025. OASIS: Open agent social interaction simulations with
one million agents. arXiv preprint arXiv:2411.11581.
(2025).
[143]
Ye, K. et al. 2019. Interpreting the rhetoric
of visual advertisements. IEEE transactions on pattern analysis and
machine intelligence. 43, 4 (2019), 1308–1323.
[144]
Ye, K. and Kovashka, A. 2018. Advise: Symbolism
and external knowledge for decoding advertisements. Proceedings of
the european conference on computer vision (ECCV) (2018),
837–855.
[145]
Zahavi, A. 1975. Mate selection—a selection for
a handicap. Journal of Theoretical Biology. 53, (1975),
205–214.
[146]
Zahavi, A. and Zahavi, A. 1997. The
handicap principle: A missing piece of Darwin’s
puzzle. Oxford University Press.
[147]
Zhang, M. et al. 2018. Equal but not the same:
Understanding the implicit relationship between persuasive images and
text. arXiv preprint arXiv:1807.08205. (2018).
[148]
Zheng, L. et al.
2024. Judging llm-as-a-judge with mt-bench and chatbot arena.
Advances in Neural Information Processing Systems. 36,
(2024).