From Stock Photos to Trillions of Artifacts: Shutterstock’s Data Journey with Jefferson Frazer

Share via Twitter Share via Facebook Share via Linkedin Share via Reddit

Get more video from Redmonk, Subscribe!

In this conversation recorded at Fastly Xcelerate, James Governor talks with Jefferson Frazer, Director of AI at Shutterstock, about how a two-decade-old media company reinvented itself for the age of foundation models. Frazer explains how Shutterstock’s long history of human-reviewed, meticulously labeled content left it uniquely positioned when hyperscalers began seeking diverse, web-scale datasets for multimodal training. Frazer digs into the technical backbone too: a 100-plus petabyte footprint, cloud-agnostic architecture built on Fastly’s private fiber and Wasm-based edge compute, and the discipline of tracking data provenance through standards like C2PA and SynthID. Looking ahead, he makes the case that unified embedding spaces and portable metadata—not any single model provider—will define the next chapter of AI-ready content.

This RedMonk video is sponsored by Fastly.

Links

Transcript

James Governor
Hey, it is James Governor, co-founder of RedMonk. We’re here for another MonkCast. We’re here with Jefferson Frazer, Director of AI at Shutterstock. And we’re gonna have a conversation where I think we kind of bridge some of the changes in business models in engendered by AI, but also really talk to the technology journey that Shutterstock has been on. So welcome to the show.

Jefferson Frazer
Thank you for having me. My pleasure.

James Governor
So yeah, I guess the first question it’s always a good one to ask. For those of us that don’t know, and I’m sure most of you do, what even is Shutterstock?

Jefferson Frazer (00:35)
Well, Shutterstock is the place that you look to for media. If you’re looking for one piece of content, or if you’re a foundational model trainer and you’re looking for that data set to push your insights to the next level, we hope that you think of Shutterstock as your provider of choice. we’ve been around for quite some time and we’ve had a very holistic focus on data sanitization and security throughout our life cycles. So we feel that we’re uniquely situated to serve both our small and ultra large customers. Okay.

James Governor (01:03)
That’s awesome. And by the way, you are really good at this stuff. That was immaculately done. So, quick question. That was very much articulated as sort of a an AI era story. When did it become obvious that it would have to be one? Like what was the what what was the moment where you sort of where it’s like, okay. The the story has to change. And we are an AI company now.

Jefferson Frazer (01:24)
That really happened right about the end of the pandemic. as the large hyperscalers started coming out with their first foundational models, we started having some really interesting conversations where they were interested not only in one part of the data set, but really the entire catalog. That they had such lofty goals for multimodal inputs and multimodal outputs, and they really needed a uniquely diverse and scaled ecosystem of content to fuel that type of insight.

And we found that over 20 years our ingestion process had really been centered on human review. So we were uniquely situated with a highly manicured, detailed, and truthful data set at the beginning of the AI wave. And we’ve been working with a lot of these hyperscalers and foundational model builders over time, and this has really helped us understand what they’re looking for: type of label generation, feature extraction, embeddings that are useful both at micro scale as well as web scale.

James Governor
Okay. And so you said that the the the the the metadata and the the the I guess the context was human created. How’s that changed in the last few years as as AI has become more part of the process?

Jefferson Frazer
So we’ve expanded our collection. We do include programmatically and agentically generated content now, but we have a strong focus on the providence of the data, right? If you feed hallucinations into your systems, you’re gonna get nothing but hallucinations out. We wanna make sure that we maintain strong ground truth against our sample sets to make sure that things like descriptions or semantic embeddings are actually depicting what’s in the content.

and by tracking which fields are generated by humans and which fields were generated by models, we can compare them over time and help add some self perpetuating feedback signals where we can see how models are improving, which models are more accurate and which ones are slightly more prone to the hallucinations we’re also fearful of.

James Governor
Okay. So in order to do all of this stuff, you’re serving obviously a huge amount of I mean I I hesitate to call it content. I must feel a bit weird. If we call it content, I I feel it doesn’t always do it justice. It’s art, it’s it’s photography, it’s it’s music as well.

Jefferson Frazer
It’s it’s everything that you could imagine. It started we started as a primarily as a media company. So music, MP3s, videos, 3D models, but we’ve really expanded over time. if it is data and it can be used to generate something that a model would require for fine-tuning, that is where we want to be and what we want to be helping to purvey. we have a strong focus on making sure that our content is easily indexable and cross-referenceable so that we have no

Crossover or bleeds between something that might be AI generated, something that might be human generated. If you’re looking for it, we want you to be able to find it no matter what it is.

James Governor
Have you got any any big numbers, any scale numbers, just in terms of what sort of volumes of of data are we talking about? What sort of network I mean we are here at a Fastly event, so I’m sure there are some questions about content serving. and I and and in fact, you know, I did hear your some of the the numbers you were giving. Like what kind of scale of operations are we talking about?

Jefferson Frazer
it’s really hard to wrap your head around the scale. The human brain isn’t super great about understanding numbers in the billions or trillions. We really do have trillions of artifacts. Our primary object stores sit at about forty petabytes. That’s what we consider our licensable customer facing data set. Our total footprint is well north of a hundred petabytes, and that includes all types of things like extracted logs, extracted features, insights that we’ve gleamed, and then the regular things that run the business.

James Governor
Okay,

that’s a lot.

What role does Fastly play in in being able to serve that amount of of data? And like I guess that’s the that’s you’ve given me the data footprint. Obviously the question about the access footprint is that’s a whole other question as well. So yeah, I’d love to get some some sense of A, what what does that look like from a network perspective? But also, yeah, so why Fastly?

Jefferson Frazer
So we choose Fastly because we like to be agnostic to the cloud providers. We need our data set to sit externally from each of the hypervisors. A lot of our customers are these hypervisors and hyperscalers, and we can’t expect them to want to run their compute wherever our object store is sitting. Additionally, the egress costs from leaving these walled gardens and these compute environments are so astronomically high that we need to prepare our data set to be as mobile as possible.

And Fastly has really shown outstandingly here. We have most of our data set sitting in a primary region behind a single pop. And that pop is connected via private fiber lines to each of the primary large cloud providers. Okay. And these dedicated fiber lines give us both the performance heuristics that we need to ensure throughput and speed, but as well as the ability just to ship mind-blowing amounts of data.

Trying to put this forty petabyte number into something that we can relate to. If you took the cell phone that I’m sure everyone has inside their pocket, took your cell phone out and think about it, if you took the newest cell phone that you could buy in the market right now, you’d need nearly thirty-five thousand of them. Thirty-five thousand brand new iPhones to hold the data set. And that’s not gonna be economical. We can’t ship shipping containers over to the data center every time you wanna plug in a workload.

James Governor
funny because I I I mean you you absolutely can’t. but a a a friend of mine and this was actually he was at EMC and they were doing the the U2 tour and the amount of data that they were generating they actually did I mean this is a few years ago they actually did take a data center around to follow the tour just because of the amount of data that was I mean he had a lot of fun on that on on you know on on that tour. But yes I am not saying that you should be doing that.

Jefferson Frazer
Data center on wheels with you two. How could that be a bad time?

James Governor
Right?

So the the the I think that the the question for me b about the or or or one of the questions about Farsi is like like so it is it is it it is providing you the ability to have that portability. on the other hand, if we think about like your needs for watermarking, all of that sort of stuff, is there

Are you using Fastly for any of the filtering? And yeah, is that is are you are you using any of the sort of that functionality g that goes beyond just cache?

Jefferson Frazer
yeah, we use a ton of Fastly features. we’re deeply embedded inside of the ecosystem, a lot of tools around observability, origin inspector. We like

The blend of compute and VCL. It’s the best of both worlds. We can have complex systems that we write ourselves or easy to manage throughput caches. We’ve spent quite a bit of time developing a hierarchical system of our compute policies so that we can reuse components. As we continue to bring on new customers, new data sets, we don’t want to have to rebuild the wheel every time. We don’t want to run these same feature extraction pipelines and rebuild them from zero. So we want to make sure that we can call out to separate components.

at our scale, it’s actually not economical for us to do just in time transcoding for videos in particular. So we try and have our data set sitting ready for whatever complexity or resolution the customer may need.

James Governor
Okay. And what do you what do you actually build in? What sort of, I mean like you the developer how many developers do you have? What sorts of environments are they using? Sure. Is that changing? Are there standards?

Jefferson Frazer
There are absolutely standards which are evolving day by day. We spent a long time developing internal software development kits and libraries, which package a lot of our proprietary routing knowledge and a lot of our feature extraction tooling. So this way we have a centralized repository. We use tools like JSII to do polygot transpiling. So we will write once in TypeScript and then output five or six times into various languages.

we feel very strongly that code needs to be locally runnable. So Fastly’s local compute execution environments is absolutely priceless for us. The ability to mimic what we’re going to see in VCL or compute locally speeds up our feedback loop so much that we can actually ship changes at high velocity with confidence in our s in our sets.

James Governor
This may sound like a dumb question, but what specifically do you mean by locally in this context?

Jefferson Frazer
On the developer’s laptop. Yeah. as the old adage goes, if we could ship the laptop we would, and that’s how Docker was born. But we love containers. we love that Fastly uses Wasm and has a portable ecosystem with repeatable compute environments.

James Governor
Yeah, I’m not sure that’s a story that the market at large has fully grocked. I feel like people are still just thinking, Fastly’s over there. Not that this is local, but one of the stories that at this event today at Fastly Xcelerate has come through pretty strongly. And now to hear it again from you, that’s increasingly important.

Jefferson Frazer
Yeah, containerization and portability is paramount. We need to be able to move not only our data sets, but also our compute to wherever we may want to run it. And Fastly gives us the flexibility to do that that we’ve never had with any other partners. Okay.

James Governor
That’s awesome. So, I mean, I I I think one of the questions for me would be like, how do you see your business? And it’s any business in the age of AI, it’s difficult to sort of really say how it will evolve. but but with the with these frontier model providers making these certainly so far quite extensive claims about how amazing their image generation will be, or music generation, or video generation.

How do you see that playing out about the sort of the balance of like the capabilities you’ve got, the deep expertise you’ve got, but there are competitive threats and new threats from new entrants that are like, no, we’ll go 100% AI.

Jefferson Frazer
Well, we’re we’re more powerful as a group. We’re more powerful with our peers. so we really see our future as being part of a network, being being part of a large practice where companies can come together to share their insights and to help partner and purvey their data. as we’ve grown over the last five years through this crazy AI expansion, we’ve really started to drift away from traditional media types and have started inspecting everything from behavioral data to audio books.

To transcoded and transpiled software libraries and templates. If if you can think of it, there’s a model out there that needs it. And we think that we’ve spent the last 20 years developing pipelines that are uniquely situated to apply labeling and best practices around embeddings and feature extraction so that we can help the entire ecosystem move forward, so that we can help our partners drive new insights into their data and help license.

And generate new revenue streams that may not have been possible before from bespoke data sets.

James Governor
really are about all kinds of data sets. And if we live in a world where, and and you know, this is potentially coming in so some geographies, where you have to show the provenance in terms of of the model that is being used, where the what materials went into that, I guess that puts you in a an interesting position.

Jefferson Frazer
It it leaves us uniquely situated to have providence data that is rich and accurate. as we’ve been so hyper focused on our data sanitization and review process throughout the years, we’re starting from such a risk rich place with such rich data practices that it only makes sense to perpetuate.

James Governor
So in order to make my my daughter happy, like how do we encourage, how do we encourage then some of the the the people building these models to start making that part of their story, to say, no, actually this is trained on, I mean, I I guess they would say, you know, shutter stock is a part of the story, but not all of the story. But how can you see a world where we begin to actually say what’s inside, you know, the the AI billing materials? Like what are the the market drivers or is is there some demand for these organisations to start saying where what the models were trained on.

Jefferson Frazer
Absolutely, and we’re starting to see regulations around the world start to catch up to where the model providers are. the two new big emerging things that we see are C2PA and SynthID. And that needs to be encrypted into the object at rest.

Can’t be separated, you can’t serve the metadata without some type of flag showing the provenance. And I really feel that it’s that’s just the natural next progression. Once you have a strong statement of certainty over your rendered output, the very next question that comes to mind is, well, where what was your input? Where did your tokens come from? Right.

James Governor
Okay. ‘Cause i we clearly have not been doing any of that thus far. it has been, you know, c you know, the the the Wild West. but but yeah, it it it it’s pretty clear that w we in order that we have people continuing to make things, to make models more interesting, we’re gonna have to find ways that support them. And I think regulation is probably gonna be part of that.

Jefferson Frazer
Yeah, and and that’s natural in any type of market system. We we see that this is moving so fast and the the cutting edge of technology is really bleeding into the future and this this type of regulation is going to be necessary going forward.

James Governor
Steam engines became safer because when they started blowing up, people said, no, we need some standards here or we do need to regulate.

Jefferson Frazer
This. Yeah, absolutely. And and hallucinations can absolutely ruin the output of a model. Okay.

James Governor
yeah, we don’t we don’t want I mean there’s enough slop already with other models just giving us pure gray goo. So okay, so technically then I guess this would be some of my last question, maybe or thereabouts. Like what are there what excites you most about this this this AI platform that you’ve got access to and you know the the the the the technology that’s coming

so quickly towards us. Like what are you excited by? Like what gets you up and you’re like, I’m a director of AI, so what comes next is this. And I’ve really got to pursue that and use that in terms of underpinning, you know, my business. Okay.

Jefferson Frazer
Unified embedding space as embedding engines are starting to mature, the number of dimensions are starting to grow. Where 125 used to be useful, we’re now seeing things come to market with 3,000 dimensions. Your rendered metadata can actually be larger than the objects that you’re putting in, which is a really interesting way to think about the problem. That I can describe something with more bytes than is actually contained in it. And

So for so long, a lot of these embedding spaces have been fragmented and segmented. If you wanted to run a search against images, you couldn’t really cross-reference that easily to a video. You could do clip detection, single frames, but it’s a process. You had to build a wheel. And now we’re seeing a very interesting space with unified embeddings where it’s becoming trivial to cross-reference content types. And this is only going to get more interesting as models start supporting more modes of input and output.

And the mind boggles at what’s possible. As we start looking at real-world models, what type of embedding space is that going to need? How many dimensions are going to be sufficient to describe that level of perspicacity and of life? That type of problem, and measuring that is really interesting to me.

James Governor
That’s awesome. And and as you say, you’ve got this you’ve got all the choice because you’ve got your data in a place, you’re you’re you’re living in Switzerland. Which of the you know, are are there hyperscalers that are doing a better job of supporting that sort of database? I mean we had that whole wave of vector databases, now it seems to be stabilizing and vectors are very much part of the data stores that we like, yeah, what’s your data store story and and what providers are you working with?

Jefferson Frazer
So we’re s we’re really seeing a a wide range of things. we’re primarily centered in Snowflake for our data inspection and curation capabilities right now. but we want to be portable. If our customers are sitting in Google and they need Gemini embeddings, we need to be able to do that for them. if they want to run vianova embeddings in AWS, we can also do that for them.

we focus both on proprietary as well as non-proprietary models for inspection into our content. So we want to make sure that we can give if you want something from the walled garden, we can absolutely help you with that. Or if you want to be absolutely agnostic and run something that’s open source, open-weighted, and we can discuss what those weights are doing to the data set, that’s absolutely within our capabilities as well. The future is metadata and being portable. Each each of the model providers and the hyperscalers.

James Governor
So the future is metadata.

Jefferson Frazer
Their tools come with trade-offs, and none of them are perfect. Each one comes with a consideration, limits in the number of fields that you can have, limits in the size of the data set that you can cross-reference. And you have to be painfully aware of those at ingestion time. And by keeping our data external from each of these hyperscalers, we can be intentional about where we want to run things. If we want to go try out a new embedding model in another cloud or try out a new vector DB, we absolutely can do that.

James Governor
That’s awesome. So I in in conclusion, I think we’re we’re lucky to have you here today. Thank you so much. But you’re lucky to be here today, aren’t you? Because tomorrow you are gonna fulfill a lifetime ambition. And I love this. So what are you what are you up to tomorrow?

Jefferson Frazer
that’s right. so tomorrow I’m going to be going to the Globe Theater to go see Much Ado About Nothing. Absolutely my favorite Shakespeare. It’s got a little of everything. We’ve got the high society humor, we’ve got the low society humor and twenty minutes about conversations on clothes. it’s got a little bit of everything and then the star crossed lovers to make it all sweet.

James Governor
Yeah, that is awesome. I am so glad that you get to see that show. And the globe is amazing. So yeah, that’s gonna be great. thanks so much for joining us. I hope you all enjoyed listening. that is a wrap from us here at Fastly Xcelerate. It’s been a MonkCast, and thanks.

Jefferson Frazer
Absolutely. Thanks for having me.

More in this series

Conversations (147)