dynamic range masterclass
We have a new video out on our YouTube channel and this one’s a real doozy. Greg dives deep into the science & art of dynamic range to answer all the questions and put this debate to rest. Enjoy!
The article below is a rough transcript of this video, which I'd appreciate you watch in its entirety.
I. Defining camera’s dynamic range
The dynamic range of a camera sensor is: the ratio of the saturated signal to the sensitivity threshold. This definition is straight from ISO 157392, the very same standard that the IMATEST software is based upon. So we have saturated signal and sensitivity threshold and the relationship between them is called dynamic range. Let’s examine what those are.
Signal saturation
Digital imaging sensor is a grid of pixels. Each pixel is a complex structure with micro lenses and color filters and electrical circuitry, but the main part is the photodiode. A photodiode is a type of diode sensitive to light, hence the name. It converts light energy into an electrical charge.
During each exposure, which let’s say for 24fps movie is 1/48th of a second, it collects photons – the light particles and turns them into electrons. It helps to think of it as a well gathering rain water. In fact it kind of looks like a well under a microscope.
So that well has a limited capacity. And when it fills up completely during the exposure it can no longer store any extra charge. That is what’s called signal saturation in our dynamic range equation. In laymen’s terms we call it clipping.
This brick wall effect is very important to understand because it is present in every and all digital sensors and it can’t be avoided or overcome. So if we are sensor engineers we can’t increase the dynamic range from the top of the well, we have to approach it from the bottom.
Sensitivity threshold
On the bottom end we have noise as a limiting factor. There are several different kinds of noise. First is shot noise which is an intrinsic randomness of light particles that can’t be avoided.
Fun fact, we imagine photons as infinitely small, but the actual number of photons hitting a single photosite during a single exposure is only measured in a few thousands. So it’s not surprising that some photosites get much more juice than others. And that randomness, that shot noise increases greatly when there is not much light to go around.
The second type of noise is dark noise which is like the self-noise of a microphone, it’s some baseline level of noise that a photosite generates just by being active.
So the shot noise and the dark noise create a baseline level of noise that the signal has to compete with. When the signal, which is light absorbed and stored as an electrical charge during the exposure, breaks through the noise floor, that is called the sensitivity threshold. In other words, it’s the minimum amount of light needed to register above the noise floor.
That point is called signal-to-noise ratio of 1 and that is what determines the bottom of our dynamic range equation. CineD uses a more strict standard of SNR2, but the the principle is the same. It’s the low end of our dynamic range equation.
SNR is key to dynamic range
So if the top part is the clipping point that we can’t change, the only thing we can do to increase dynamic range is to somehow lower the sensitivity threshold. And we can either go about it by reducing the effects of shot noise or by reducing dark noise.
Shot noise can be mitigated by making the photosites bigger. Because all else being equal, a bigger bucket collects more rain per exposure time than a small bucket. That’s why the most sensitive cameras have low megapixel counts. Alexa famously has only 7.5MP in open gate, the reigning dynamic range champion Alexa 35 has only 14.5MP and my trusty FX6 from Sony has 10MP in its biggest video format.
Apart from using bigger pixels you can also work on improving the sensor architecture. For example reducing the wasted space between pixels, or putting the circuitry behind the photosite. That’s why backside-illuminated (BSI) sensors are much more sensitive than their FSI predecessors.
Reducing unwanted reflections and light scattering inside the sensor chamber also improves things. That’s why when designing Alexa 35 ARRI had to come up with new bayonets because the old ones weren’t dark enough for such a sensitive sensor.
So that’s what you can do about the shot noise. Bigger buckets and less wasted photons. As for dark noise, one major thing you can do is to keep the sensor at a stable temperature. Noise has direct relation to heat. Digital sensors need to operate at a stable temperature around 40C to produce the least amount of noise. That’s why the bulk of a proper cinema camera is occupied by a massive heatsink. And Alexas famously have a Peltier element in them so they can dynamically cool or heat the sensor keeping it stable even in incredibly harsh weather conditions. By the way that is why ARRI cameras don’t require black balancing.
There’s a bunch more super complicated things one can do to reduce the self-noise of a sensor but these are the big ticket items: big unobstructed photosites sitting in a perfect dark at a comfy stable temperature.
DR is about shadows, not highlights
Before we descent deeper into the rabbit hole, I want to call your attention to the backwards logic of the DR equation. When we discuss dynamic range as cinematographers or just camera nerds, all we talk about are the highlights. That’s where the artistic value is for us, protecting the highlights. High dynamic range means not blowing out a window in a day interior shot or improving the look of light sources and specular highlights, right? Wrong! That’s not where dynamic range comes from.
In the realities of sensor design, dynamic range is all about seeing into the shadows. It’s all about improving the sensitivity of sensor, because that’s the only part that CAN be improved. You can’t raise the clipping point, once you blew out a window, it’s gone. But if you stop down on exposure to protect the window, a more sensitive camera will allow you to still see into the shadows. And later in post production you can re-map where middle grey was supposed to be and return to a “normal” exposure while gently rolling off the hightlights.
That’s how dynamic range really works. And it is somewhat counterintuitive to how we’re thinking about DR. Stew on that for a while, because that’s going to be hugely important in the next part.
II. ADC: the dynamic range bottleneck
Ok, so far we’ve been living in the analogue world of photons and electrons, light and voltages. But to record an image a digital camera needs to convert that voltage into digital values. And that’s where things get a whole lot more complicated.
All that we’ve discussed so far establishes the outer boundaries of dynamic range possible for a given sensor. Higher sensitivity sensors are capable of greater dynamic range. But then why doesn’t my FX3 that is super sensitive can’t capture more than 12 stops? Well, that’s because it is limited by a 12-bit analogue-to-digital converter or ADC. The sensor is the same, but the ADC now is the limiting factor.
Let’s go back to the buckets of water analogy. You have a bucket and it collected some amount of rain water. But you need to digitize that amount to record it in your notebook. Let’s take a measuring stick and put 10 notches on it. It’s not gonna be very precise, now, is it? So let’s cut 10 more increments between each notch, and now it has a 100.
Theoretically there’s no limit to how precise you can measure it. But as armchair sensor engineers we have to decide, how precise can we afford to be. Because the more precise the measurement the more processing power, the more memory bandwidth and the bigger thermal envelope it’s gonna take. A camera can’t cost a billion dollars, require a hydro-electric plant to power it and weigh a 1,000 lbs.
So when you decide on the bit-depth of your analogue to digital converter, you put a hard cap on the dynamic range you can capture. Most video cameras in say $2,000-15,000 range have 12-bit ADCs and thus physically can’t resolve more than 12 clean stops of dynamic range.
So that statement alone makes the whole DR discussion suddenly way less interesting, but let’s examine it further.
14-bit photo mode VS 12-bit video mode
I’m one of those crazy people that takes photos on their FX3, just because that’s the only photo-capable camera I have. And every time I sit down to edit my photos in Lightroom, I’m amazed at how much dynamic range there is in those RAW files and how much colour information there is. And I realize I’m never this happy with my color grades for video. I can get it into a decent shape but I’m never as proud and happy as I can be with my photos from the same camera. Why is that?
Well, it’s because like with most hybrid cameras, the FX3’s sensor operates in 14-bit readout mode when shooting photos. What that means is that there are 16,384 notches on the measuring stick for each little bucket. 16,384 possible values from clipping point to the very bottom of the noise floor. When shooting video the FX3’s sensor switches into 12-bit readout mode, which is only 4,096 possible values, exactly four times less. After that there’s debayering, and compression into a 10-bit container, which is only 1,024 values, again four times less.
The further downstream we go in the imaging pipeline, the more DR we can lose. But we can never gain it back or get more than we had on the previous step. A 12-bit ADC camera can’t record more than 12 stops of dynamic range, period.
Linear VS Logarithmic
At this point you probably might be confused with all this talk about bit depth. Isn’t that related to the color? And why are bits equated to stops of dynamic range, aren’t those apples and oranges?
First of all, forget about the colors. We are only concerned with dynamic range here, which refers only to the range of luma values. Color appears way later in the imaging chain at the debayering stage. And we are discussing ADCs here, which do their thing even before RAW. So we are only talking about light. A camera is just a light-gathering device, as great Steve Yedlin says.
The problem with light is, is that it’s linear and our perception of it is not. And digital cameras are fundamentally linear too. But what the hell does it mean? Well, linear means that there’s 1:1 relationship between light and information. Think of it that way: each little light particle carries with it a fixed amount of information. More particles – more information.
But we humans don’t see light that way. For us to register a change of exposure of 1 stop, the light need to double or halve its intensity. This is called a logarithmic scale. Bits are a logarithmic scale too, just like light, so 1 bit is doubling or halving of the values. These are all powers of two. That’s why we use mathematical tool to record stops of light, because it is appropriately logarithmic.
So back in the linear world, it means that every stop of exposure generates double the electric charge than the previous one. And this is the craziest thing. Remember those 16,384 possible values in a 14-bit ADC? Well, half of those values are occupied by one stop, the brightest stop before clipping. And half of the remaining half if occupied by the second stop. And half of that quarter is the third stop. And so on. You get it.
That’s how linear processing works. It is true to how light works but it’s terribly inefficient for the purpose of human vision. Because we are most sensitive to the midtones and shadows because that’s where our lives happen. No information that is essential to our survival as a species is ever in the highlights. Otherwise we’d be staring into the sun all day. If anything, we avoid bright highlights, squinting and turning away and wearing polarized sunglasses.
Magic of log: bending the stick
So how’s that all relevant to the dynamic range discussion? We’ve established that Dynamic range comes from the shadows. And now that we’ve examined the linear nature of camera sensors, it turns out there’s very very limited space left for shadows and midtones on our measuring stick. Because the bulk of the well is occupied by highlight information that is of little value to us.
That’s why we need that 4-fold increase in the amount of code values to discern just 2 more stops of light in the bottom there. Because all of that granularity is needed at the bottom end of the range. And that is why the bit-depth of the ADC puts a hard cap on the potential top-end dynamic range of a sensor.
Linear data is great for cameras and computers to crunch numbers but is next to useless to us humans. So the camera packages that information it in a more efficient way, which is called a logarithmic gamma. That’s what our visual system does, it expands the shadows and compresses the highlights, so that we can pay more attention to subtle contrast differences in the midtones.
On a chart like this, linear looks logarithmic and log looks linear, because it evens out the amount of code values dedicated to each stop
It’s funny and counterintuitive but if you put linear and logarithmic gamma on a chart, it’s the log that looks linear and the linear that looks logarithmic. But it’s easier to think of it as bending the stick.The camera takes our linear measuring stick and bends it so it can fit into a smaller logarithmic container. In the process it discards a lot of the excess highlight information and redistributes the values so that every stops occupies roughly the same amount of bits inside the container. That’s how it can fit 12 stops of dynamic range into a 10-bit file.
But the important number is the bit-depth at which the initial linear processing happens. You might have heard how Sony likes to boast about their 16-bit linear raw output of its cameras, when in actuality it is recorded as 12-bit log ProRes RAW. That’s just one in the long line of marketing tricks that work only because we are stupid monkeys that start salivating when we see bigger numbers.
They might as well be outputting a 16-bit linear signal through the HDMI, but that doesn’t mean a thing. It’s like putting a huge pipe after a small pipe. You won’t get more water out of it.The ADC still reads out at 12 bits and that’s the maximum amount of stops you can get from that camera.
The only way to expand that dynamic range without increasing the bit-depth of the initial readout is something like dual gain architecture. ARRI Alexa famously has a pair of ADC streams at different amplification levels – one at zero gain, AKA the native ISO, and one at negative gain. Each ADC is 14-bit readout, but then the best parts of two images are combined into one within a 16-bit linear space. The result is then encoded as 12-bit log ARRI RAW. Alexa 35 ups that to two 16-bit readouts processed in 18-bit linear space and recorded as 13-bit log ARRI RAW. The bit-depth of the final container is always just big enough to contain the whole of the log curve, so there’s no wasted bits.
III. Cutting through the marketing BS
So, dynamic range numbers. How come there’s so much contention and confusion about them? If what I said about the bit-depth of ADC is true, there should really be no space for debates and different interpretations of dynamic range.
Now, it’s going to sound like I’m criticizing the work that CineD and Gerald Undone are doing, but I’m not. I think overall they are doing way more good than harm for the industry and definitely serve the community in a positive way. However, their overemphasis on IMATEST results and lack of nuanced discussion often creates as much confusion and misinformation as it’s trying to alleviate.
Just quoting the IMATEST numbers is like only reading the abstracts for scientific studies. Not only is it not the full picture, it can oftentimes lead you to wrong conclusions. If you really want to understand what’s going on you have to inspect both the XYLA chart images and the camera in question.
IMATEST numbers can easily be tricked and inflated by several things: in-camera noise reduction, downsampling, compression, highlight recovery and loss of contrast.
In-camera noise reduction
First let’s tackle in-camera noise reduction, because it’s the biggest culprit. This can be easily spotted as suspiciously clean and flat noise floor on the XYLA-21 chart. And it is mostly present in mirrorless hybrid-type bodies where you can’t turn off noise reduction. If a sensor is trapped in a tiny weather-sealed mirrorless body, you can be damn sure it’s hot and noisy in there.
Noise reduction makes the image cleaner at the expense of detail and color fidelity, but it doesn’t actually improve the light-gathering capability of the camera, which is what we are trying to measure.
Anything that was done to the noise after the RAW stage does not reflect the true SNR, and thus the dynamic range, of a sensor. Sure it improves some arbitrary number that we like to argue about, but in actuality it is shadow clipping. Whatever details that were buried in the noise floor are now dead and gone forever.
So generally the more pristine the noise you can record the more you can recover from it. This is the only reason why RAW has a potential for greater dynamic range. Because it’s a way to circumvent the internal noise reduction. With advanced pre-debayer noise reduction in post you can dig out at least one good stop if not more. That is all of course depending on the “rawness” of RAW in question.
Before we move on to the next trick, there’s an exception we need to mention and that is DGO – dual gain output. DGO sensors, like ARRI Alexa and some of Canon’s cameras have cleaner shadows which is noticeable both on the Xyla chart and in normal use. DGO usually results in an honest increase of dynamic range of about 1 stop. It’s definitely a nice to have feature but it comes at the expense of slower readout. Which is a tradeoff we are going to see again and again.
Downsampling or pixel binning?
Downsampling to a lower resolution can also increase or inflate the IMATEST score, depending on what is truly going on in there. True supersampling must be 4-to-1, so 8K to 4K like on Kinefinity Mavo EDGE or something akin to what Blackmagic Design is doing with their URSA 12K RGBW sensor. Downsampling yields an increase of dynamic range of about half as stop, so not much really.
But with other cameras you have to dig deeper into what exactly is going on there. For high megapixel mirrorless bodies it’s usually some kind of weird pixel-binning line-skipping mish-mash, which is not true downsampling. It’s just a way for a camera to do less work, because it doesn’t have the processing power nor the cooling capabilities to deal with so many pixels.
codec Compression and bitrate
High-compression codec like h.265 combined with insufficient bitrates also effectively obfuscates the noise by averaging pixels and throwing away a bunch of fine detail information in the process. Image quality suffers, but IMATEST score improves.
These first three effects – noise reduction, downsampling and compression – are often compounding and found together on high-megapixel mirrorless bodies. And that’s the main reason why you see some clearly 12-bit ADC cameras like Panasonic S1H or Sony A7 IV score 13 stops on IMATEST.
algorithmic Highlight recovery
The next gimmick that can easily fool IMATEST is highlight recovery. This refers to the reconstruction of clipped highlights when debayering RAW sensor data. This is the trick that RED relies upon to vastly inflate their dynamic range numbers. 20 stops, anyone?
The idea here is that pixels under different color filters clip at different levels. Knowing that you can extrapolate to recover some of the lost color information in the highlights. But it is of course just guesswork and not real light gathered by the camera. The feature can work nicely sometimes, but other times it produces awful plasticky-looking highlights or weird artifacts.
I think most people would agree that those are fake stops and you can’t count them towards the camera’s true dynamic range.
Loss of contrast
The last thing that can inflate the dynamic range numbers is loss of contrast due to poor optics or poor sensor housing design. We are either talking about a sensor chamber that poorly deals with reflections. Or we have a bottleneck at the lens level that can’t produce sufficient contrast. But usually sane people don’t use vintage lenses for dynamic range tests, so this is seldom an issue.
The loss of contrast is easy to spot because the shadow stops on the Xyla chart will be too closely spaced, almost overlapping each other. And the blacks will appear milky, as if you have a black pro mist filter on the lens.
We need to do better
As you can see, much more care needs to be taken so we don’t compare apples to oranges. For example, you can’t directly compare a RAW-shooting cinema camera with highlight reconstruction to a weather-sealed high-megapixel mirrorless hybrid that shoots H.265 with tons of internal noise reduction. The IMATEST scores in this example are next to useless as points of comparison. And I won’t even mention throwing phones into the mix, with all their computational trickery.
I can pretty accurately ballpark a camera’s dynamic range just by looking at it and considering its sensor specs, form-factor, power draw, and price. The range is not that big anyway, it’s bound to be somewhere between 11 and 14 stops at SNR=2. Anything higher than that is an Alexa LF or 35.
The problem is that camera companies usually don’t disclose the particulars of their sensor design and image processing pipeline. Especially if it would reveal some dirty tricks that they’re not proud of. So there’s bound to be some guessing or speculation by our budding scientists. Which leaves even more space for confusion and misinterpretation.
But I guess in the spirit of journalistic integrity, you have to ask those questions to the camera companies. And then the onus is on them to be transparent and honest or cagey and secretive about it. There’s a fine line between revealing your trade secrets and substantiating your outrageous marketing claims.
And if RED’s example has thought us anything, is that people are resourceful and will eventually always find out the truth. And that the loss of reputation for BSing your customers is permanent.
IV. The price of high dynamic range
If we were to discuss specific cameras, I’m sad to report that, much like in other areas of life, you generally get what you pay for.
Like I said many times already, the vast majority of affordable prosumer and professional cameras on the market are limited by 12-bit readout in video mode, some are even 11-bit. The only notable exception would be the Fujifilm X-H2S which boasts BSI stacked sensor capable of a 14-bit readout in real-time framerates. But that camera is so full of compromises that it only serves as an exception that proves the rule.
The other honorable mentions are Canon’s DGO cameras which eek out an extra stop, putting them slightly above the 12-bit crown in terms of dynamic rage. Canon C500II is most likely a 14-bit camera, considering its original price and the sensor inherited from C700. And you can find used well under $10K.
But these are outliers. Generally speaking, if you want to jump from the 12-stop world to 14-stop world, you are looking at a huge price hike. Prepare to spend at least fifteen grand on a used Amira or more likely twenty-twenty five grand for a used Alexa Mini if you’re into anamorphic lenses.
V-Raptor is somewhere in there dynamic range-wise, although with all the needed accessories the realistic price is more in the 30 grand zone. And the new global shutter sensor is bound decrease dynamic range somewhat.
Alexa Mini LF is a side grade to the full frame look, so the extra half a stop of DR is hardly worth mentioning. It’s still a current camera, and considering the original mini doesn’t meet the 4K mandate, Mini LFs are still very much in demand and going for 60-70 grand used.
Up from that is Alexa 35 which will set you back roughly a hundred K for a ready-to-shoot package. Funnily enough camera prices seem to track with the logarithmic nature of light. Each extra stop of dynamic range is a doubling of the price.
And that is to be expected because every extra stop has a design and manufacturing cost associated with it and I’m not only talking about cash.
Sensor design is one huge tradeoff
The camera design is one huge balancing act, you increase one spec at the cost of another. Camera technology is amazing and always improving, but there’s no circumventing the laws of physics, something can’t come out of nothing. A 14-bit or higher ADC readout takes longer, requires more power, more memory bandwidth, higher-speed storage and more of it.
For reasons we discussed in chapter one, dynamic range comes at the expense of resolution and framerates. A highly sensitive camera can’t be high resolution, that’s why we only see about 12-13 clean stops from Sony Venice II, even though it’s most likely a 16-bit ADC camera. Same for the little Burano.
But Venice 2 has the highest readout speed on the market, the 8K makes it versatile with multiple format support, it has decently high speeds, higher base ISO and of course the Rialto system. All that makes it a better choice for certain productions, better than Alexa 35. Because highest dynamic range or best overall image quality might not be a primary concern.
Nobody wants a garbage image of course, but that’s not really an issue these days. You can get a great-looking image from almost any decent video-centric camera these days. I mean, Gareth Edwards moved from an Alexa 65 on Rogue One to the FX3 on the Creator. Clearly the 12 stops of dynamic rage and the color fidelity of ProRes RAW image out of the FX3 were enough for a Hollywood blockbuster movie.
Cinema-ready image quality is a low-hanging fruit these days
That choice was dictated almost entirely by the 12,800 ISO and the tiny form-factor of the camera. Both of these had cascading effects on the whole production – smaller camera supports, smaller lights, smaller crews. That downsizing was desired by the director to move faster, be more responsive and ultimately create more freely. The price of the camera had nothing to do with it.
Dynamic range is not the end-all be-all
The higher you set that image quality bar for yourself, the more trade-offs and sacrifices you’re gonna have to make. Which leads us to conclusion that prioritizing dynamic range is a conscious choice in camera design. And the reason we don’t see more high dynamic range cameras is because for a lot of applications, that extra dynamic range is just not worth the tradeoff.
12-bit ADC cameras produce image quality that is good enough for the vast majority of applications, while offering a number of benefits that the big-boy cameras don’t have. ARRI cameras have the highest dynamic range because they only make cameras for one application – cinema. So the tradeoff of resolution, framerates and the increased production overhead in terms of power and data is worth it to them. It’s right there in the mission statement for the Alexa: to produce a camera with the best overall image quality.
Different jobs call for different tools
That is a conscious design choice, but it doesn’t mean it’s the only right one. Camera is a tool and just like with any craft, different jobs require different tools. Dynamic range is not the one ring to rule them all. For some applications other camera specs are more important:
You need higher resolution for multiple format support and some VFX applications,
You need higher framerates for sports, wildlife and product cinematography,
You need a higher base ISO, lower power draw and more efficient codecs for documentary applications.
The ultimate best image quality is just not the top priority for a lot of applications. In fact, one could argue, that Alexa 35 is already a step into the placebo territory.
You can’t handle the best IQ
I think for a lot of us camera nerds image quality is an engaging theoretical exercise, and not a real practical concern. I mean, wouldn’t it be cool to film your TikToks with 17 stops of dynamic range in 13-bit uncompressed RAW?
We think we want the highest image quality, the highest dynamic range, but I’m here to tell you – you can’t handle it. That supreme image quality comes at a cost. I’m not even talking about the price of the camera package. If Elon Musk showed up at your door right now and handed you an Alexa 35, you wouldn’t know what to do with it.
Do you know that that camera sucks down a one hundred and fifty watt-hour brick every hour and requires a bespoke 24 volt battery solution? And do you know that it generates 2TB of RAW data every hour? You wouldn’t have the batteries for it, you wouldn’t have the media for it, hell, you probably wouldn’t even have a proper tripod for it.
And you sure as shit wouldn’t have storage space and upload bandwidth to deal with that footage. We are talking about an enterprise-level NAS just to store it, not to mention some kind of enterprise cloud solution for backups and transfers.
You wouldn’t know what to do with it, you’d be absolutely paralysed by that camera. Just think about that for a second. With the best camera in the world you would be just sitting there scratching your head, being less creative than you could be with just the phone in your pocket.
V. How much dynamic range do we really need?
And that brings us to the real question that sparked all of this. How much dynamic do we really need and why are we so fixated on it?
Well, that last part is easy. Anybody who ever shot a video in their life encountered the problem. You are shooting a day interior. If you expose for the inside, the windows are blown out. If you expose for the windows, the inside is in total darkness. That frustrating experience is at the core of the desire for more dynamic range. The logical and natural conclusion is that we need the camera to see what our eyes can see. But I would put forth that is a fool’s errand and I’ll explain why later.
Now consider that a modern smartphone, utilising an exposure bracketing technique and some computational trickery can produce an image that solves our problem. Both the inside and the outside are properly exposed for. And yet, we look at that image as cinematographers as image makers and is it desirable? Does it look better than what our 12-stop professional cameras can capture?
Of course not, it looks like dogshit – all plasticky and fake. So what’s the deal here? What do we really want? Are we looking to hit a certain number? Or are we after a certain kind of look? Well, I’ll separate my answer into two parts: technical and artistic.
Technical side
From the technical perspective we need to understand the concept of mastering and where all that dynamic range is going to end up.
You are watching the video a top of this page in SDR, which is 8-bit and thus cannot display more than 8 stops of dynamic range. 8-bit is just 256 values per channel which is 28, so 8 stops. Why so low and why doesn’t it look all that bad?
The reason for this is efficiency. As we’ve established, humans as a species are acutely sensitive to midtones. If we reference Ansel Adams’s Zone System, the important stuff in our lives happens in zones 2 to 7, so only about 6 stops. These stops in SDR are linear, meaning information in them tracks more or less 1 to 1 with how exposure really changes. The rest gets heavily compressed, deep shadows occupy the lower half a stop basement while all the highlights are squeezed into the top 1,5 stops.
And contrary to popular misconseptions, HDR doesn’t actually change anything in the midtones. It’s right there in the Davinci Resolve guidelines for SDR to HDR conversion. Everything under 50 IRE remains the same, stuff between 50 and 90 IRE gets very slightly expanded while highlights in the 90 – 100 zone get greatly expanded. HDR is only about opening up those compressed highlight stops, which increases the scene’s overall contrast ratio leading to a punchier, more vivid appearance. But those 6 stops that hold the midtones and shadows pretty much stay exactly where they were in SDR.
So really the only technical reason to seek higher dynamic range is the appearance of highlights. And that all has to do with that brick wall effect that is the curse of all digital sensors. Once the signal is saturated, all the detail is gone. And that transition into overexposure is abrupt and produces a very undesirable appearance especially when brought down in post.
Film vs digital
Film is often perceived as the gold standard for dynamic range but in reality it was never that great. Moreover, it’s hard to quantify exactly due to the non-linear nature of the medium. For example, Kodak Vision 3 200T film stock resolves only about 10 stops linearly, but at the extremes there’s that built-in analogue roll-off. That natural-looking loss of saturation and detail makes it very hard if not impossible to detect the clipping point.
Through a stroke of luck or a genius, I have no idea, but film is much better at mimicking our own visual experience than digital. And it has dictated the look of cinema for almost a 100 years. But it’s not about the number of stops, what is really being praised here is the appearance of highlights.
Now with modern post-production workflows, digital cameras can replicate the highlight response of film. Meaning, as long as there’s no hard clipping, a roll-off curve can be applied in post-production to compress the highlights without losing detail. Even some light clipping is acceptable and can be further massaged with techniques such as halation, blooming, low contrast filters or vintage lenses or adding grain in post-production. All these can soften the transition into the lost highlights, as long as the clipped areas are not too large. But the general principle of digital cinematography is to expose for the highlights that you can’t control.
Scene DR vs camera DR
So it is part of your job description as a cinematographer to control and compress the dynamic range of the scene. That means exposing the scene for the camera and not the camera for the scene. In other words, you have to bring in some damn lights.
Lighting has a huge part to play in filmmaking and obviously there’s an enormous artistic aspect to it that we will discuss later. But just from the technical side, you need to expose for what you can’t control and then use light to reduce the dynamic range of the scene.
Back to our day interior example it might lighting the inside with a powerful daylight fixture. Depending on the outside conditions you might need something like an Aputure 600D or stronger to match the sun’s intensity. If we are talking about a huge space you might need 18K HMIs with a whole crew to run them, but the idea is the same.
And before we start whining about not having big enough lights, I want to stress that there’s always something you can do. First of all, you can work the windows themselves: you can net them, which will be invisible when out of focus, you can put ND film on them, you can hang some sheers and curtains, if there aren’t any.
Even if you have zero budget and zero time, there’s always some measure of control you have just by choosing your angle and shot size. Just put the light source out of the frame for starters.
If Roger Deakins blows out a window, you can be sure it's intentional
Sometimes even a simple reflector is enough to wrap that light around your subject. OR – embrace it and play the whole thing in a silhouette. OR – a shocker – let the windows blow out, that’s a look. Just make sure they’re blown out completely and throw on a black pro mist to diffuse the edge. Check out this scene from Queen’s gambit. That’s a creative choice, not an accident or a mistake.
It’s not the size and budget of your production, it’s your willingness to do the necessary work to make it look halfway decent. Right now in 2024, your camera’s Dynamic range is just not a big enough technical concern unless you are shooting on a potato. And this is a hard pill to swallow because it means you can’t buy your way to good cinematography.
Squeeze & drop technique
What a high dynamic range camera really allows you to do is postpone exposure decision into post production. In other words, it allows you to be sloppy and make mistakes.
If you download some sample footage from ARRI and plug it into Resolve, after a color space transform you will see that the dynamic range of the camera far exceeds the SDR space. So far in fact, that you can’t use this footage as is, you have to compress it somehow to avoid clipping.
And as you start color grading this footage what you’ll naturally arrive at is something called Squeeze and Drop technique. It can be done a lot of different ways but the principle is the same, you lower the exposure until the highlights sit right, then you raise the black point until you restore all the necessary shadow detail. Then you finesse the midtone contrast, which is the most important part.
The end result is that the important stuff is compressed into the bottom 60% of the waveform, preserving all that dense shadow detail while only letting the highlights peak into the 70s, 80s and 90s, depending on their brightness. What I just described is essentially the exposure technique that needs to be done on set, but postponed into postproduction.
Of course exposure latitude is great and useful and it bears stating that modern cinematography has one foot firmly planted in postproduction. Sometimes it takes 30 minutes to achieve on set what can be done in two minutes in post, and you must know when to make that trade-off. But generally speaking, the higher the level of production is the closer the gap is between what was captured on set and what was done in post.
VI. Numbers VS Art
Realistic VS cinematic
What I’m trying to say here is whatever great dynamic range you capture, you will end up squeezing it all in essentially a 6-stop range with some room for highlight roll-off.
And HDR doesn’t actually change this. Yes it expands highlights, but the important part of the image stays the same. The goal post hasn’t moved, it’s not a total rewrite of the rule book. HDR doesn’t mean that we are going to need sunglasses to watch outdoor scenes. We are not trying to blind the audience.
The idea that we want the camera to see what our eyes see is ultimately a fool’s errand because movies are not life and life is not movies. I mean, I don’t know about you, but my life sure doesn’t look like a Roger Deakins movie. The lighting sucks, the set design is boring, the wardrobe is atrocious and there’s no makeup to speak of. Plus everybody forgets their lines all the time and the plot doesn’t move unless you do stuff.
What I’m saying is, movies don’t look like life, cause that was never the point. Movies look like movies, that’s what “cinematic” means. It is motion captured at 24 frames a second, at 180 degree shutter and displayed on a flat surface. Every single attempt to alter that formula in recent history has failed.
People don’t want realistic. People want cinematic. At least for now. Cinema is a relatively young art form, it can and most likely will change. Maybe in 50 years movies will be a totally virtual experience, who knows.
But right now, cameras are not eyes and eyes are not cameras and the whole notion of realistic capture is fundamentally misguided.
A 20-stop camera would hinder this shot, not help it
In fact, even if we could capture a 20-stop scene it would look flat and boring on screen. It would lack the contrast, the richness and density that make a cinematic image appealing in the first place.
It’s not about the camera
Let’s look at some stills from my favorite DPs. What’s the dynamic range of these shots? 4, maybe 6 stops at most?
I look at this frame and ask myself – could this have been captured with my camera? Yes, easily. Could I capture it? Well, I could imitate it. But to conceive of it in the first place? And to create it on set on the day, under pressure in a high-stress time-starved environment? No, I’d be too busy puking from the anxiety.
It’s not about the camera, it’s not even about having a good eye. I have a good eye, every Tom, Dick and Harry has a good eye, it’s not that rare and not that valuable. What’s rare and valuable is having decades of managerial and technical experience translating your good eye into actual pictures on screen.
Rembrandt had 6 stops of dynamic range
Let’s go even further and consider the art of painting.
If you would bother to bring a spot meter into a museum, you’d be surprised to find out that Rembrandt has like 6 stops of dynamic range. Just one frame if that is worth millions of dollars. I have twice as much stops in my camera, shooting up to 120 frames per second, yet how much are they worth?
The idea that we can reduce art to some simple number or a metric is ridiculous and yet, in our space, it’s ubiquitous. Because the underlying hope is that by increasing some kind of magic number we somehow can elevate our craft into the higher planes. If only we had greater dynamic range, if only we had more pixels, more frames per second, bigger sensors.
A painter is not concerned with the quantity of strands of hair in their paintbrush. They are only concerned with what it lets them paint. Yet we spend valuable time arguing about pixels in the comment section. If numbers were truly important, then the camera with the biggest numbers would be the choice of greatest of filmmakers, right?
RED is nowhere to be seen, despite touting 8K 120fps specs
But just one look at Oscar nominations for cinematography this year tells a different story. RED is the company build on the American ethos of making the numbers bigger, yet, where are they? Or look at Sundance this year, it is dominated by the Alexa Mini. An old camera that is not Netflix approved, has 3.4K resolution at best and caps out at 30 FPS in OG RAW. And yet, all these independent filmmakers choose this camera of their own volition. Because there’s much more to a creative tool that cannot be reduced to numbers.
Alexa is still a go-to choice for independent filmmakers
Despite its meagre specs on paper
Art is about less, not more
But I would go further and throw out even the qualitative metrics. The very idea that you need more of something to create is just a form of procrastination. And I’m talking to myself right now as much as anybody else.
The whole notion that “if only I had more X” is a lullaby that we sing to ourselves to cope with that fact that we’re not moving towards our dreams. Online camera debates are a comfortable distraction that numbs the existential dread of meaningful work that we know we should be doing.
The fact that I haven’t made my feature film yet, is not because I don’t have an Alexa. It’s because I haven’t sat down to write it. Instead, I spend countless hours researching the minutiae of cameras I have no business even considering. The idea of more is a poison to creativity.
Every artist worth their salt is on a quest of downsizing. It’s always about less, not more. All the best shots have two, maximum three colors. Some have one. Dynamic range is tiny. Compositions. All seemingly so simple.
It’s the amateur who asks for more. The pro asks for less. Malevich tried to reduce it to one color, one shape. Jack White is trying to reduce it to one string. Roger Deakins is trying to reduce it to one light. Everybody who hears the muse is trying to chisel down this granite of reality down to its core, to its truest nugget.
Because they are on the quest for the essence, the truth. It’s a hard path to walk, and it takes decades, lifetimes even, but the work of a master is unmistakable as it transcends time itself.
The whole point of this essay is to convey one idea, that I’m sure you already know in your heart. It’s not. About. The camera.
Conclusions
It would be a cop-out to end this without talking actual numbers. After all, the question still stands — how much dynamic range do we actually need? With all of that’s been said, from both the technical and artistic perspective in mind, here’s my tier list of dynamic range.
10 stops
10 stops is the bare minimum for professional applications. Great care must be taken to reduce the dynamic range of the scene to avoid clipping. And that makes it less suitable for uncontrollable lighting conditions. That being said, there have been feature-length movies shot on 10-stop HD cameras that launched directing careers. So there are no excuses.
12 stops
12 stops is enough for most applications. Some acceptable clipping is expected in the brightest highlights (like light sources, specular reflections) which can be further massaged with filtration or in post. This is most likely the camera you have right now and much like Ken, it’s Kenough for whatever artistic endeavor you can envision.
14 stops
14 stops is on the top end of analogue film capabilities and is enough to cover most challenging scenes and vastly improve the appearance of highlights, providing sufficient data for filmic roll-off in post. This is the Alexa level, the camera that brough digital acquisition over the hump of early adoption and that’s been dominating the filmmaking landscape ever since.
16 stops
16+ stops surpasses the DR of film and approaches the limits of human vision. This is Alexa 35 and it is the new industry standard for high-end productions. Her huge dynamic range is used to further improve the appearance of super bright and saturated highlights, like neon lights in the night or emergency beacons and brake signals on vehicles.
But for the vast majority of scenes, the dynamic range difference between Alexa and Alexa 35 will be imperceivable to the end viewer.
18 stops
As of the time of this writing 18+ stops only exist in RED’s marketing materials. This is placebo territory, where gains in image quality will be so minuscule and come with such a heavy cost that it might not be worth it to go there. But I’ve been wrong before, so who knows?
Will we see 14-stop cameras under $10K?
Is it likely we will see video-centric cameras with 14 stops under $10K in the near future? Yes, with BSI Stacked sensors and 14-bit readout, it’s probably feasible with current tech. But the resolutions and frame rates will probably take a hit, which the head honchos don’t like. Big numbers are easier to sell.
We should probably expect the next dynamic range leap from Canon. They have had next-gen DGO sensors in development for years now. And they have nothing to lose, no Venice 2 to protect and their market share is steadily decreasing. So it could be a tremendous comeback for them.
But will it suddenly make your stuff look 4 times better? No. 10% better? Yes, maybe. For the rest you’ll have to work on what’s in front of the camera.
The irony of dynamic range
In conclusion I can’t help but to point out the irony – the more you control the lighting the less DR you actually need. A high DR camera allows more freedom for run-and-gun filmmaking. Because you can half-ass the exposure and still recover everything in post.
But the landscape of camera market is such that that high DR is only available to the least run-and-gun of productions. It might be hard to admit, but you can’t run-and-gun your way to great cinematography. At some point you actually have to stop and think about the image you are trying to create. And that’s when you’ll see what’s the real limiting factor is.
It’s not the camera you have or don’t have.
It’s you.
