Holly Huey.

Applied Scientist @ Adobe | Experimental Psychology, PhD | Building GenAI image & video model evals to enhance human creativity

what I do

I am the Technical Lead for Adobe Firefly Image & Vector Eval for 1P & 3P models. I build human eval tools & auto evaluators for model quality assessments that are benchmarked on specialized, large-scale eval sets. To curate these sets, I hire globally and culturally diverse teams of professional photographers & editors, graphic designers, and video creators & editors to custom create, curate, and & annotate assets. My research statistically measures improvements & regressions in model behavior related to: text-to-visual & visual-to-visual systems, usability, and editability

what my impact is

My evaluations determine whether 1P models are ready for public release and guide leadership decisions for buying 3P models & weights. To do this, I develop production-readiness standards based on large-scale user research and co-create PRDs with product managers for model quality. I also collaborate with companies like Google & OpenAI to evaluate their pre-release models before Day 0 launches. My work is also deeply integrated with AI Safety and IP Guardrails teams to ensure that models do not create or enable the creation of harmful content.

who I am

I have been passionate about visual media throughout my entire life, truly. I grew up on Whidbey Island, WA (a very rural island), and I spent my teen years sketching and hiding away in art studios. I did 2 art apprenticeships and then double majored in Philosophy and History of Mathematics & Sciences from St. John's College, where I was fascinated by how novel visualizations have catalyzed & been used to communicate humanity's major scientific discoveries to the global public (e.g., Cartesian coordinate system, periodic table, Vitruvian man) —  yet, despite this critical role of visual media in human innovation, relatively little is known about how people communicate in visual form. This question continues to inspire me, academically, professionally, and personally.

I earned my PhD in Experimental Psychology in the Cognitive Tools Lab at UC San Diego, studying human and model generative behaviors. My dissertation research evaluated how user goals shift visualization strategies and downstream interpretation by others. I studied this by generating & analyzing large-scale datasets of drawings, diagrams, and data visualizations. In other words, I studied how people make pictures. Today, I study how GenAI models make pictures.

Selected GenAI research highlights

image generation & editing

At Adobe, I evaluate text-to-image, image-to-image editing, vectorization and image-to-video editing via instructional prompting, masking, layer segmentation, and more. This work spans both multi-modal foundational and specialist models. By measuring SOTA model behavior via human & auto evals, I can help executive leadership and product managers understand where the market is today and define the market of tomorrow. I also work closely with model engineers to help direct model development and ensure alignment of real user use cases.

cross-cultural image models

My research supports the development of culturally-specific GenAI models. I hire and closely work with global teams of artists with extensive lived experience from specific cultures and countries. This human-in-the-loop collaboration ensures that our models do not contribute to erasure or inappropriate Westernization of non-Western cultures. My work has contributed to text-to-image models built in the Adobe x HUMAIN global strategic partnership to build GenAI Arabic models.

HUMAIN Image 1 generated image of Elephant Rock with a hot air balloon HUMAIN Image 1 generation settings with a falconer and falcon HUMAIN Image 1 generated ad layouts with Arabic and English text
video editing

Occasionally, I also conduct mixed-methods qualitative user research when data access or interpretation is limited. I strongly believe that user research keeps GenAI research honest and grounded. I have interviewed filmmakers, video editors, and musicians to develop GenAI video editing parameters based on users' creative workflows. These insights helped frame large-scale annotation experiments that eventually fine-tuned different model parameters. For example, our Creativity & Cognition 2024 paper evaluated how people (N >800) and LLM models select B-Roll to enrich videos depending on their goals to make entertaining or informative content.

sketching

During my PhD, I specialized in large-scale human eval methods. Back then in 2019-2024, large-scale crowdsourcing techniques for human eval were novel — this is now the bread & butter (or more strongly, the backbone) of GenAI model evals & post-training that everyone conducts, but I am proud to have been among the first to develop these tools during grad school and take this early expertise into my current professional research career. In my dissertation work, I used these human eval methods to curate large-scale datasets of sketches and human annotations (e.g., >90K human & GenAI sketches). Datasets like these can be used for a wide array of research domains such as computer vision, attention, memory, semantic segmentation of objects, and information prioritization based on user goals. See my Publications below to access those datasets, such as our Nature Communications 2024 paper on evaluating concept development in children through their drawings.