The Data Moats Hiding in Nature: Building a Competitive Edge with Proprietary Data

Insights
Anna Bosch
September 22, 2026
min read

The Data Moats Hiding in Nature: Building a Competitive Edge with Proprietary Data

Insights
Anna Bosch
Published on
Sep 29, 2026
min read
Back to stories

The Data Moats Hiding in Nature: Building a Competitive Edge with Proprietary Data

The Data Moats Hiding in Nature: Building a Competitive Edge with Proprietary Data

Insights
Anna Bosch
September 22, 2026
min read

‍

Proprietary data: moat in times of AI

‍

In the ongoing debate about what constitutes a defensible moat in the age of AI, proprietary data has emerged as one of the most compelling answers. Yet one category of proprietary data has received surprisingly little attention: data generated by nature.

‍

Nature is constantly producing vast amounts of information – across ecosystems, soil, water, weather, crops, forests and the atmosphere. But unlike much of the data AI companies work with today, this information is not readily available to scrape, license or retrieve through an API. Much of it first has to be observed and collected.

‍

This makes nature data particularly interesting from a venture perspective. Its inherent characteristics – physicality, hyperlocality and constant change – make it difficult to collect at scale, but for the same reasons potentially difficult to replicate. 

‍

As the cost of collecting and processing physical-world data comes down, we believe a new class of companies can turn this challenge into a powerful competitive advantage.

Why nature data matters now

In addition, nature data itself is becoming more important across various industries. According to the World Economic Forum, more than half of global GDP is moderately or highly dependent on nature and the services it provides, from pollination over water regulation to soil stability. Two economic factors drive that increasing importance:

‍

  • Differentiated inputs become more valuable
    In the past years, AI models have turned from competitive advantage into an easily accessible commodity. Today, the moat is not the model a company uses, but what it uses to feed that model. If two companies can use similar models, but one has accumulated five years of proprietary atmospheric, hydrological or biological observations that the other cannot access, the data generation engine becomes the true moat. This is particularly true when the data is tied to the physical world. A competitor may reproduce software surprisingly quickly, but it cannot retroactively deploy sensors across thousands of sites and recreate five years of historical observations.

  • Physical-world information is more mission-critical than ever
    A rapidly changing climate, vulnerable infrastructure, and geopolitical conflicts have made nature data increasingly important to high-stakes decision-making. Agriculture needs better soil and weather information to manage inputs and predict yields. Insurers and financial institutions need more granular physical-risk data. Industrial companies need environmental intelligence around physical assets. The same underlying capabilities are increasingly relevant to defense and security: our portfolio company LiveEO, for example, is expanding its satellite- and AI-powered geospatial intelligence platform from critical infrastructure monitoring into dual-use defense applications. In addition, growing regulatory requirements add another source of demand, particularly around biodiversity, water and land use. 

‍

Alongside the growing importance of nature data, the technological development of the past years has led to a dramatic decrease in cost when it comes to data collection. Sensors that power data-collection systems have become smaller, cheaper, and more capable, while the means of connectivity have improved. More and more satellites provide increasingly granular sensing data, and autonomous systems like drones open up entirely new data collection methods. AI makes it possible to process data volumes that would be prohibitively expensive to interpret manually. All of these developments lead to a reality where large amounts of data can be continuously collected at increasingly shrinking costs.

‍

The companies chasing the nature data flywheel

Building on these developments, we have seen a couple of companies in the past years that leverage nature data while owning and expanding data-generation infrastructure.

‍

  • NatureMetrics uses environmental DNA to generate species-level biodiversity data. 
  • Pivotal collects primary biodiversity observations and builds a data layer on top. 
  • Sorcerer uses autonomous weather balloons to collect atmospheric data. 
  • Hula Earth combines in-situ sensors and bioacoustics with satellite imagery. 
  • WaterSense deploys autonomous sensor stations that deliver water quality data.
  • Perennial and Seqana are using digital soil mapping and satellite-based measurement to generate soil organic carbon data.

‍

At first glance, these companies operate in different markets. But the underlying architecture is remarkably similar: A well-designed proprietary collection infrastructure delivers unique, highly valuable physical-world data. Having that data collection and data quality advantage allows these companies to develop better models and insights compared to their competitors, which in turn attracts more customers, motivating even more deployments. An increase in deployments directly feeds again into the quality of the proprietary data collection infrastructure. 

‍

That vicious cycle is what we coin the nature data flywheel. If that loop works, the valuable asset is not necessarily the sensor, the model or the dashboard. It is the data-generation engine underneath them.

‍

Looking ahead: key determinants for the future

We are not yet convinced every nature data company earns the "moat" label just because it collects proprietary data. Three questions separate the durable businesses from the interesting pilots.

  1. Does the data get better with scale, or does it just get bigger?
    The strongest datasets should become more predictive, more granular or more valuable as coverage increases. Simply accumulating terabytes is not enough.

  2. Who pays, and how reliably?
    Regulatory tailwinds are powerful, but they are also a single point of dependency. The strongest companies in this space will find commercial buyers who need the data for operational decisions, not only for compliance filings.

  3. Can the collection economics scale?
    Physical-world data has an inherent disadvantage relative to software: collecting it costs money. Sensors break, devices need maintenance, and samples need processing. The best businesses therefore need declining marginal collection costs or rapidly increasing value per unit of collected data.

There is also an additional question that may prove particularly important: does time itself become part of the moat? A company with a ten-year longitudinal dataset may have an advantage that a new entrant literally cannot accelerate its way into. Those are the data businesses we find particularly interesting.

Summary

As we have seen, the rise of AI has caused competitive moats to shrink significantly. Nature data turned out to be an ideal proprietary moat, as it is hard to collect, yet hard to commoditize. The key determinant for success, however, is not just owning a high-quality proprietary dataset. The companies that will thrive in the upcoming years will be those that manage to build a scalable, resilient, and efficient data collection engine that improves with every deployment – enabling software or models that become better as that dataset expands. 

‍

‍

Proprietary data: moat in times of AI

‍

In the ongoing debate about what constitutes a defensible moat in the age of AI, proprietary data has emerged as one of the most compelling answers. Yet one category of proprietary data has received surprisingly little attention: data generated by nature.

‍

Nature is constantly producing vast amounts of information – across ecosystems, soil, water, weather, crops, forests and the atmosphere. But unlike much of the data AI companies work with today, this information is not readily available to scrape, license or retrieve through an API. Much of it first has to be observed and collected.

‍

This makes nature data particularly interesting from a venture perspective. Its inherent characteristics – physicality, hyperlocality and constant change – make it difficult to collect at scale, but for the same reasons potentially difficult to replicate. 

‍

As the cost of collecting and processing physical-world data comes down, we believe a new class of companies can turn this challenge into a powerful competitive advantage.

Why nature data matters now

In addition, nature data itself is becoming more important across various industries. According to the World Economic Forum, more than half of global GDP is moderately or highly dependent on nature and the services it provides, from pollination over water regulation to soil stability. Two economic factors drive that increasing importance:

‍

  • Differentiated inputs become more valuable
    In the past years, AI models have turned from competitive advantage into an easily accessible commodity. Today, the moat is not the model a company uses, but what it uses to feed that model. If two companies can use similar models, but one has accumulated five years of proprietary atmospheric, hydrological or biological observations that the other cannot access, the data generation engine becomes the true moat. This is particularly true when the data is tied to the physical world. A competitor may reproduce software surprisingly quickly, but it cannot retroactively deploy sensors across thousands of sites and recreate five years of historical observations.

  • Physical-world information is more mission-critical than ever
    A rapidly changing climate, vulnerable infrastructure, and geopolitical conflicts have made nature data increasingly important to high-stakes decision-making. Agriculture needs better soil and weather information to manage inputs and predict yields. Insurers and financial institutions need more granular physical-risk data. Industrial companies need environmental intelligence around physical assets. The same underlying capabilities are increasingly relevant to defense and security: our portfolio company LiveEO, for example, is expanding its satellite- and AI-powered geospatial intelligence platform from critical infrastructure monitoring into dual-use defense applications. In addition, growing regulatory requirements add another source of demand, particularly around biodiversity, water and land use. 

‍

Alongside the growing importance of nature data, the technological development of the past years has led to a dramatic decrease in cost when it comes to data collection. Sensors that power data-collection systems have become smaller, cheaper, and more capable, while the means of connectivity have improved. More and more satellites provide increasingly granular sensing data, and autonomous systems like drones open up entirely new data collection methods. AI makes it possible to process data volumes that would be prohibitively expensive to interpret manually. All of these developments lead to a reality where large amounts of data can be continuously collected at increasingly shrinking costs.

‍

The companies chasing the nature data flywheel

Building on these developments, we have seen a couple of companies in the past years that leverage nature data while owning and expanding data-generation infrastructure.

‍

  • NatureMetrics uses environmental DNA to generate species-level biodiversity data. 
  • Pivotal collects primary biodiversity observations and builds a data layer on top. 
  • Sorcerer uses autonomous weather balloons to collect atmospheric data. 
  • Hula Earth combines in-situ sensors and bioacoustics with satellite imagery. 
  • WaterSense deploys autonomous sensor stations that deliver water quality data.
  • Perennial and Seqana are using digital soil mapping and satellite-based measurement to generate soil organic carbon data.

‍

At first glance, these companies operate in different markets. But the underlying architecture is remarkably similar: A well-designed proprietary collection infrastructure delivers unique, highly valuable physical-world data. Having that data collection and data quality advantage allows these companies to develop better models and insights compared to their competitors, which in turn attracts more customers, motivating even more deployments. An increase in deployments directly feeds again into the quality of the proprietary data collection infrastructure. 

‍

That vicious cycle is what we coin the nature data flywheel. If that loop works, the valuable asset is not necessarily the sensor, the model or the dashboard. It is the data-generation engine underneath them.

‍

Looking ahead: key determinants for the future

We are not yet convinced every nature data company earns the "moat" label just because it collects proprietary data. Three questions separate the durable businesses from the interesting pilots.

  1. Does the data get better with scale, or does it just get bigger?
    The strongest datasets should become more predictive, more granular or more valuable as coverage increases. Simply accumulating terabytes is not enough.

  2. Who pays, and how reliably?
    Regulatory tailwinds are powerful, but they are also a single point of dependency. The strongest companies in this space will find commercial buyers who need the data for operational decisions, not only for compliance filings.

  3. Can the collection economics scale?
    Physical-world data has an inherent disadvantage relative to software: collecting it costs money. Sensors break, devices need maintenance, and samples need processing. The best businesses therefore need declining marginal collection costs or rapidly increasing value per unit of collected data.

There is also an additional question that may prove particularly important: does time itself become part of the moat? A company with a ten-year longitudinal dataset may have an advantage that a new entrant literally cannot accelerate its way into. Those are the data businesses we find particularly interesting.

Summary

As we have seen, the rise of AI has caused competitive moats to shrink significantly. Nature data turned out to be an ideal proprietary moat, as it is hard to collect, yet hard to commoditize. The key determinant for success, however, is not just owning a high-quality proprietary dataset. The companies that will thrive in the upcoming years will be those that manage to build a scalable, resilient, and efficient data collection engine that improves with every deployment – enabling software or models that become better as that dataset expands. 

‍

Go to website
URL copied to clipboard

Learn the Essentials of Entrepreneurship

Discover our curated collection of tools, best practices and relevant articles. Get started now.
Explore our startup resources