Holly Huey.

Applied Scientist @ Adobe | Experimental Psychology, PhD | Improving GenAI image & video models to enhance human creativity

What I do

I am the Technical Lead for Adobe Firefly Image & Vector Eval for 1P & 3P models. I build human eval tools & auto evaluators for model quality assessments benchmarked on specialized, large-scale eval sets. To curate these sets, I hire globally and culturally diverse teams of professional photographers & editors, graphic designers, and video creators/editors to create & annotate assets. My research statistically measures improvements & regressions in model behavior related to: text-to-visual & visual-to-visual systems, usability, and editability

What my impact is

My evaluations determine whether 1P models are ready for public release and guide leadership decisions for buying 3P models & weights. I also work with companies like Google & OpenAI to evaluate their pre-release models before Day 0 launches. Collaborating with AI Safety and IP Guardrails teams is deeply integrated into my work to ensure that models do not create or enable the creation of harmful content.

Who I am

I have been passionate about visual media throughout my entire life, truly. I grew up on Whidbey Island, WA (a very rural island), and I spent my teen years sketching and hiding away in art studios. I did 2 art apprenticeships and then double majored in Philosophy and History of Mathematics & Sciences from St. John's College, where I was fascinated by how novel visualizations have catalyzed & communicated our most major scientific discoveries (e.g., Cartesian coordinate system, periodic table, Vitruvian man) —  yet, despite this critical role of visual media in human innovation, relatively little is known about how people communicate in visual form. This question continues to inspire me, academically, professionally, and personally.

I earned my PhD in Experimental Psychology in the Cognitive Tools Lab at UC San Diego, studying human and model generative behaviors. My dissertation research evaluated how user goals shift visualization strategies and downstream interpretation by others. I studied this by generating & analyzing large-scale datasets of drawings, diagrams, and data visualizations. In other words, I studied how people make pictures. Today, I study how GenAI models make pictures.

GenAI research highlights

image generation & editing

At Adobe, I evaluate text-to-image, image-to-image, image-to-video, and image editing via instructional prompting, masking, layer segmentation, and more. By measuring SOTA model behavior via human & auto evals, I can help product managers and exec leadership understand where the market is today and define the market of tomorrow.

cross-cultural image models

My research supports the development of culturally-specific GenAI models. I hire and closely work with global teams of artists with extensive lived experience from specific cultures and countries. This human-in-the-loop collaboration ensures that our models do not contribute to erasure or inappropriate Westernization of non-Western cultures. My work has contributed to text-to-image models built in the Adobe x HUMAIN global strategic partnership to build GenAI Arabic models.

HUMAIN Image 1 generated image of Elephant Rock with a hot air balloon HUMAIN Image 1 generation settings with a falconer and falcon HUMAIN Image 1 generated ad layouts with Arabic and English text
video editing

Occasionally, I also conduct mixed-methods qualitative user research. I strongly believe that user research keeps GenAI research honest and grounded. I have interviewed filmmakers, video editors, and musicians to develop GenAI video editing parameters based on users' creative workflows. These insights helped frame large-scale annotation experiments that eventually fine-tuned different model parameters. For example, our Creativity & Cognition 2024 paper evaluated how people (N >800) and LLM models select B-Roll to enrich videos depending on their goals to make entertaining or informative content.

sketching

During my PhD, I specialized in large-scale human eval methods for collecting and analyzing sketches. Datasets like these can be used for a wide array of research domains such as computer vision, attention, memory, semantic segmentation of objects, and information prioritization based on user goals. See my Publications below to access those datasets, such as our Nature Communications 2024 paper on evaluating concept development in children through their drawings.