RevoData https://revodata.nl your Databricks partner Fri, 07 Aug 2026 10:01:48 +0000 en-US hourly 1 https://wordpress.org/?v=7.0.3 https://revodata.nl/wp-content/uploads/cropped-Layer-1-1-32x32.png RevoData https://revodata.nl 32 32 What is Databricks? https://revodata.nl/what-is-databricks/ Fri, 07 Aug 2026 08:54:35 +0000 https://revodata.nl/?p=7724

Databricks is a cloud-based data and AI platform. It helps organizations collect, prepare, govern, analyze, and use data for reporting, machine learning, AI applications, and operational decision-making. For decision-makers, the simplest explanation is this: Databricks gives data teams one shared environment to work with large amounts of data and turn that data into trusted analytics and AI. Instead of maintaining separate platforms for data engineering, data warehousing, data science, machine learning, and governance, Databricks brings those capabilities together on one open foundation. That matters because many organizations have the same problem: data is spread across systems, teams use different tools, AI initiatives struggle with data quality, and reporting definitions are inconsistent. Databricks is designed to reduce that fragmentation.

Want to understand whether Databricks matches your organization? RevoData can help you define a focused proof of concept with clear business value, technical scope, and success criteria.

What is Databricks?

Databricks is a unified platform for data, analytics, and AI. It is used by data engineers, analysts, data scientists, machine learning engineers, AI engineers, and business teams that need reliable insights from large and complex datasets. The platform started as an open-source data engineering ecosystem. Databricks was founded in 2013 by people behind major open-source data technologies, including Apache Spark, Delta Lake, and MLflow. Today, the platform has expanded beyond big data processing into data warehousing, governance, real-time processing, AI development, application development, and data sharing.

A practical way to understand Databricks is to compare it with a factory for data and AI:

  • Raw data comes in from applications, files, databases, sensors, APIs, and cloud storage

  • Data engineers clean, structure, and combine it

  • Governance controls who can access which data

  • Analysts use it for dashboards and SQL reporting

  • Data scientists build models

  • AI teams build assistants, agents, or prediction systems

  • Business users consume the results through dashboards, applications, or automated processes

A pipeline tool and a governance tool sitting next to each other don’t help much if someone still has to manually hand off results between them. The real value comes from connecting data engineering, analytics, AI, and governance into one operating model, so a dataset built once can move straight into a dashboard, a model, or an application without a manual handoff in between.

What problems does Databricks solve?

Most organizations do not have a shortage of data. They have a shortage of usable, trusted, and well-governed data. Common symptoms of this include:

  • Reports that show different numbers for the same metric

  • Data teams spending too much time on manual preparation

  • AI pilots that cannot move into production

  • Cloud costs that are hard to explain

  • Data stored in separate systems with unclear ownership

  • Slow access to new datasets

  • Governance rules that are inconsistent across tools

Databricks addresses these issues by creating a shared data foundation. It combines the flexibility of a data lake with the reliability and performance expected from a data warehouse. This architecture is often called a lakehouse. In plain language, a lakehouse allows organizations to store large volumes of different data types while still applying structure, quality controls, and governance. That makes it suitable for both traditional analytics and AI workloads.

Who is Databricks for?

Databricks is mainly intended for organizations that need to work with data at scale. It is especially relevant when data is strategic to the business and when multiple teams need to collaborate on the same data foundation.

Typical users include:

  • Data engineers. They use Databricks to build pipelines, ingest data, transform datasets, automate workflows, and prepare reliable data products.

  • Data analysts. They use Databricks for SQL analytics, dashboards, and reporting on governed datasets.

  • Data scientists. They use Databricks to explore data, train models, run experiments, and collaborate with engineering teams.

  • Machine learning and AI engineers. They use Databricks to build, evaluate, deploy, and monitor models, AI applications, and agents.

  • Platform and governance teams. They use Databricks to manage access, lineage, quality, cost controls, and workspace standards.

  • Business leaders. They do not usually work in Databricks every day; however, they benefit from faster analytics, better AI readiness, and more reliable decision-making.

Databricks is less relevant when an organization only needs small-scale reporting from a single application or when a simple spreadsheet-based process is still sufficient. It becomes more valuable when data volume, complexity, governance needs, or AI ambitions increase.

What does Databricks do exactly?

Databricks supports several core functions.

1. Data engineering

Data engineering is the process of moving, cleaning, and preparing data. Databricks helps teams build pipelines that can handle batch and streaming data. For example, an organization can ingest transaction data every hour, combine it with customer data, validate the results, and publish a clean dataset for reporting. For decision-makers, this means fewer manual exports, fewer fragile scripts, and more repeatable data processes.

2. Data warehousing and BI

Databricks can also be used for SQL analytics and business intelligence. Teams can query curated datasets, build dashboards, and serve reporting tools with governed data. This is relevant for organizations that want one platform for both data engineering and reporting, rather than moving data through multiple separate systems.

3. Machine learning and AI

Databricks supports the full machine learning lifecycle: experimentation, feature preparation, model training, deployment, and monitoring. It also supports generative AI use cases, such as retrieval-augmented generation, AI assistants, and domain-specific agents. The key point for decision-makers is that AI quality depends heavily on data quality. Databricks helps connect AI development to governed enterprise data instead of isolated prototypes.

4. Governance and security

Databricks includes governance capabilities that help teams manage access, metadata, lineage, and policies across data and AI assets. This is critical when data includes customer information, financial data, operational data, or sensitive business logic.

A strong governance layer helps organizations answer questions such as:

  • Who can access this dataset?

  • Where did this number come from?

  • Which dashboards use this table?

  • Which models depend on this data?

  • Is this data approved for AI use?

5. Real-time and streaming use cases

Some organizations need to act on data quickly. Examples include fraud signals, sensor monitoring, logistics events, customer behavior, security telemetry, or operational alerts. Databricks supports streaming data workflows, which means teams can process data as it arrives instead of waiting for a daily batch.

6. Data and AI applications

Databricks is increasingly used to build applications that work directly with enterprise data. These may include internal tools, AI assistants, planning applications, or operational dashboards. This is important because many organizations want AI to move beyond experiments. They need applications that are secure, governed, and connected to current business data.

Databricks on Azure, AWS, and GCP

Databricks runs natively on three major clouds, with the same core platform, Spark, Delta Lake, notebooks, MLflow, SQL analytics, underneath each one. What differs between them is the integration layer: how Databricks connects to each cloud’s identity system, storage, networking, and native services. Most organizations choose based on where they already run infrastructure, not on which cloud runs Databricks best.

Databricks on Azure

Many Dutch organizations use Databricks through Azure Databricks. Azure Databricks is the Databricks platform integrated with Microsoft Azure services, identity, networking, and cloud infrastructure. It holds first-party status on Azure, meaning support cases can route through Microsoft’s own enterprise support channels and billing can run through an existing Azure Enterprise Agreement. For organizations already committed to Azure, especially ones also running Power BI or Microsoft Purview, this can mean one bill, one identity system, and one support relationship instead of three.

Databricks on AWS

AWS was the first cloud Databricks launched on, and it remains the most established of the three deployments. Databricks on AWS integrates with Amazon S3 for storage, AWS IAM for identity and access management, and AWS PrivateLink for private network connectivity, alongside complementary services like Redshift and AWS Glue. Organizations that already run their core infrastructure on AWS often choose Databricks on AWS for exactly that reason: it slots directly into an environment they’ve already built, rather than asking them to adopt a second cloud’s identity and storage model.

Databricks on GCP

Databricks on Google Cloud is the newest of the three deployments, running on Google Compute Engine and integrating with Google Cloud Identity, Google Cloud Storage, and BigQuery. The standout feature here is BigQuery federation: teams can query Delta Lake tables in Databricks alongside BigQuery datasets without moving data between the two. That matters most for organizations already invested in BigQuery for SQL analytics that want to add Databricks for Spark workloads and ML training without replacing what’s already working. Vertex AI integration connects models trained in Databricks to Google’s own inference infrastructure.

Choosing between them

For organizations in Amsterdam and the wider Netherlands, or anywhere else, the decision is often less about the brand of cloud and more about the operating model. Questions to ask include:

  • Which cloud platform do we already use?

  • Where is our data stored?

  • What are our security and compliance requirements?

  • Which teams need access?

  • Do we have the skills to run Databricks effectively?

  • Which first use case can prove value quickly?

RevoData helps organizations answer those questions and translate them into an implementation path, on whichever of the three clouds that turns out to be. The skills question deserves a direct answer: if the goal is running Databricks well without building that operational capability in-house first, Managed Databricks lets RevoData run the platform on your behalf, handling day-to-day operations, security, and reliability while your team focuses on using the platform rather than maintaining it.

Databricks compared with alternatives

Choosing a cloud is one part of the decision. The other part is knowing how Databricks itself stacks up against other platforms you might be weighing instead. Decision-makers often compare Databricks with other analytics platforms, cloud-native data warehouses, or integrated BI and data environments. The right choice depends on the organization’s data maturity, cloud strategy, and use cases.

A warehouse-first platform can be a strong fit for structured reporting, SQL workloads, and dashboarding. An integrated analytics suite can be attractive for organizations that want tight alignment with office productivity, BI, and low-code tooling. A specialized machine learning platform may fit teams focused on model development.

Databricks is often strongest when the organization needs one open foundation for data engineering, analytics, AI, and machine learning at scale. It is especially relevant when data is stored in a lakehouse architecture, when teams work with large or mixed data types, or when AI initiatives need governed access to enterprise data.

A useful decision rule:

  • Choose a BI-first tool when the main need is business reporting

  • Choose a warehouse-first approach when the main need is structured analytics

  • Choose Databricks when the need spans data engineering, analytics, governance, machine learning, and AI

  • Choose a combined architecture when different teams need different interfaces on top of the same governed data foundation

RevoData helps organizations avoid tool-led decisions. The better starting point is the business outcome, the data landscape, and the operating model.

Databricks certification and skills

Choosing Databricks is one decision. Having the skills to actually run it well is a separate one, and that’s what certification is meant to support. Databricks certification helps professionals prove practical knowledge of the platform. Certifications are available for roles such as data engineer, machine learning practitioner, and solution architect.

For organizations, certification matters because Databricks projects require more than access to the platform. Teams need to understand architecture, governance, cost management, pipeline design, data modeling, security, and deployment patterns.

RevoData’s strength is that its consultants are Databricks-certified, and RevoData itself is a Gold-certified Databricks consulting partner. This means common architecture mistakes and expensive rework get caught during design instead of after go-live, which helps clients move faster from assessment to delivery. RevoData also has a strong focus on continuous learning and has one of the highest numbers of Databricks Champions in EMEA.

Databricks Community Edition and Free Edition

Certified expertise has to start somewhere, and for a lot of people, that starting point is trying Databricks hands-on before committing to anything larger. For years, Databricks Community Edition was a common way for individuals to try Databricks. That has now been replaced by Databricks Free Edition.

For decision-makers, the important point is that a free environment is useful for learning the concepts, but it isn’t a substitute for a properly scoped implementation, one with real governance, representative data, cost visibility, and clear success criteria behind it. A demo environment can show that the platform works. It can’t tell you whether it solves your actual business problem.

Current valuation and market position

Once you’ve tried the platform itself, the next natural question is usually about the company behind it: is this a safe long-term bet? Databricks is one of the most visible private companies in the data and AI market. Its valuation and revenue growth show that the platform has significant market traction, especially as organizations invest in AI and governed data foundations.

For buyers, valuation should not be the main decision factor. It does, however, indicate that Databricks is not a niche tool. It is a major enterprise platform with a large ecosystem, strong investor backing, and continued product investment. What matters more than either of those signals is fit: whether Databricks lines up with your use cases, existing architecture, skills, governance requirements, and expected business value.

How to start with Databricks: a PoC path via RevoData

A successful Databricks proof of concept should be focused, not trying to rebuild the entire data platform at once. RevoData typically recommends a practical PoC path for those seriously considering the platform:

  1. Select a high-value use case. Choose a use case with measurable business value. Examples include faster reporting, improved forecasting, customer segmentation, operational analytics, geospatial analysis, AI readiness, or automated data quality checks.

  2. Assess the current data landscape. Map the relevant source systems, data owners, quality issues, access rules, and existing reporting flows. This prevents the PoC from becoming a technology demo without organizational context.

  3. Define success criteria. Success criteria should be specific. Examples include reducing processing time, improving data freshness, replacing manual steps, improving governance visibility, or enabling a model to move into production.

  4. Build a minimum viable architecture. Create a focused Databricks setup with the required ingestion, transformation, governance, and consumption layers. Keep the scope narrow enough to deliver evidence quickly.

  5. Evaluate business and technical results. At the end of the PoC, decision-makers should know what worked, what needs improvement, what skills are required, and what a broader rollout would involve.

  6. Create a roadmap. A PoC should lead to a roadmap. That roadmap may include platform standards, data product design, governance rollout, cost management, training, and migration priorities, along with a decision on whether the platform will be run in-house or through Managed Databricks.

Why work with RevoData?

Following that process on your own is entirely possible, but the speed and risk profile change considerably with the right partner. RevoData helps organizations turn Databricks from a platform choice into a working capability.

As a Databricks Gold Partner with 100% Databricks-certified consultants, RevoData combines technical expertise with delivery experience. The team helps clients define use cases, design architecture, build production-ready pipelines, set up governance, and transfer knowledge to internal teams. A PoC is part of that broader Databricks consultancy. Once a platform is live, that same team can hand off to Managed Databricks for day-to-day operation, so the people who built the platform aren’t also the ones stuck maintaining it indefinitely.

This is especially valuable for decision-makers who want to reduce risk. Databricks can deliver significant value, but only when implementation choices match the organization’s goals, skills, and data maturity.

Ready to explore Databricks with a focused proof of concept? RevoData can help you define the use case, architecture, and success criteria before you scale.

Final thought

Databricks is best understood as a shared foundation, one platform that engineers, analysts, data scientists, and AI teams all build on, rather than a tool that belongs to just one of them. That’s what turns enterprise data into analytics and AI at scale: reliable, governed, reusable data products everyone can build from.

For decision-makers, that reframes the real question. It’s less about whether to buy Databricks and more about which business problem to prove first, and what architecture you’ll need to scale from there. RevoData helps answer that question, with certified Databricks expertise, practical implementation experience, and a clear path from PoC to production, whether that ends with an internal team running the platform or RevoData doing it through Managed Databricks.

FAQ's

Azure Databricks is Databricks integrated with Microsoft Azure. It allows organizations to use Databricks within an Azure environment, including Azure identity, storage, networking, and security patterns. It also holds first-party status on Azure, so support and billing can run through an existing Azure relationship rather than a separate one.

Databricks on AWS is Databricks integrated with Amazon Web Services, using Amazon S3 for storage, AWS IAM for identity and access management, and AWS’s own networking and security tools. It’s the original cloud Databricks launched on, and it’s often the natural choice for organizations that already run their infrastructure on AWS.

Databricks on GCP is Databricks integrated with Google Cloud, running on Google Compute Engine and using Google Cloud Identity, Google Cloud Storage, and BigQuery. Its standout capability is BigQuery federation, querying Delta Lake tables in Databricks alongside BigQuery data without moving anything between the two.

Use Databricks when your organization needs to process large or complex datasets, build reliable data pipelines, combine analytics with AI, govern data centrally, or support machine learning at scale. It is especially useful when data engineering, analytics, and AI teams need to work on the same foundation.

Databricks runs in the cloud and provides workspaces where teams can ingest data, build pipelines, run SQL queries, train models, manage governance, and serve outputs to dashboards, applications, or AI systems. It uses scalable compute and a lakehouse architecture to support both analytics and AI workloads.

No. Technical teams usually build and manage the platform, but business teams benefit from faster reporting, better data quality, and more reliable AI outcomes. Business users may consume Databricks outputs through dashboards, applications, or AI assistants.

Yes, as long as the PoC is focused. A good PoC should test a real business question, use representative data, and include clear success criteria. RevoData can help design a PoC that gives decision-makers evidence for the next investment step.

Yes. Managed Databricks is a dedicated RevoData service that takes over day-to-day platform operation: security patching, cost monitoring, performance, and reliability. It’s built for teams that want the platform running well without building that operational capability internally first.

]]>
Choosing the Right Geospatial Solutions for Your Data Platform https://revodata.nl/choosing-the-right-geospatial-solutions-for-your-data-platform/ Fri, 07 Aug 2026 08:30:25 +0000 https://revodata.nl/?p=7719

Pick the wrong geospatial setup, and you’ll find out the hard way that the GIS platform your team trusted for years can’t keep up once GPS data starts arriving by the million, or the scalable data platform someone built can’t tell a valid polygon from a corrupted one because nobody with spatial expertise was in the room. Either mistake costs a rebuild, and rebuilds are expensive in ways the original decision never imagined they could be.

Location data has moved out of specialist GIS departments and into the wider data platform. Mobility, customer demand, asset risk, logistics, network coverage, public space, climate exposure, service accessibility: all of it touches location now, and all of it eventually raises the same practical question: Which geospatial solutions should you choose to handle your ambitions?

This article gives you a clear way to answer that, based on your workload and users, not on whichever vendor got to you first. The short version: for most mature organizations, the answer is a combination of GIS tools for expert workflows with Databricks as the scalable geospatial backbone. What follows is how to know exactly which pieces you need and when.

Want to connect GIS expertise with scalable analytics? Explore RevoData’s geospatial solutions on Databricks.

What makes a geospatial solution actually work

A strong geospatial approach is the practical use of geographic information, data engineering, and analytics to create better location-based decisions that reduce costs, speed up operations, or support other concrete business goals. It combines geospatial technology, business data, and repeatable data processes to turn geospatial data into something an organization can actually act on.

In this context, doing this well goes well beyond making a map: it’s the ability to turn geographic information into operational insight. Examples include assigning customers to service areas, measuring travel-time access, identifying asset exposure, analyzing mobility patterns shaped by human activity, or enriching BI dashboards with regional context.

A mature geospatial solution often includes several layers:

  • source systems that contain addresses, coordinates, boundaries, or movement data

  • GIS tools for spatial experts

  • data pipelines for processing and quality checks

  • a governed data platform for storage, transformation, and access control

  • mapping, BI, or application layers for business users

  • machine learning or AI models that use spatial features

Location data that stays siloed rarely gets reused, resulting in the same address list getting cleaned three separate times by three different teams, and nobody’s sure which version is current. Geographic information should be available as a reusable data product across the organization, validated and maintained once, not rebuilt every time someone needs it.

Types of geospatial solutions

There are several categories of geospatial solutions. Each category solves a different problem.

1. GIS tools

A geographic information system (GIS) is built for spatial professionals. These tools support map creation, geometry editing, projection management, spatial analysis, layer management, and visual inspection.

Choose a GIS tool when users need to:

  • create or edit spatial layers

  • inspect maps visually

  • work with parcels, networks, regions, or boundaries

  • validate geospatial information manually

  • publish map layers for specialist users

  • perform detailed cartographic work

These tools are mostly used by urban planners, environmental scientists, utility engineers, and government agencies tracking anything from land parcels to the impact of climate change on a region.

GIS tools are valuable because spatial data is complex. Coordinate systems, geometry validity, topology, scale, and interpretation all require expertise to maintain accuracy. A data platform does not remove that need. However, GIS tools are not always the best place for large-scale analytical workloads. If a team needs to process millions or billions of location records, combine those records with CRM or sensor data, or refresh outputs automatically, a data-platform approach is usually more suitable.

2. Mapping and GeoApps

Mapping solutions and GeoApps focus on communication and user interaction. They make spatial information accessible to non-specialists through dashboards, web maps, portals, or embedded application features.

Choose mapping software when the main goal is to help users answer questions visually, such as:

  • Which areas need attention?

  • Where are service gaps?

  • Which locations are high risk?

  • Where are customers, assets, or incidents concentrated?

  • How do regional patterns change over time?

A well-designed map can make complex data easier to understand. But if the map carries the full burden of data preparation too, errors and inconsistencies end up baked into a dashboard that people trust at face value simply because it looks polished. For reliable decision-making, the layers behind the map need clear definitions, quality checks, and governance, not just a clean interface.

3. Spatial databases and geospatial engines

Spatial databases and processing engines are used to store and query geometries at scale. They support functions such as distance calculations, containment checks, intersections, buffers, and spatial joins.

These tools are useful when spatial logic needs to be repeatable. Instead of manually drawing conclusions from maps, teams can express the logic in SQL, Python, or automated pipelines.

This category becomes important when organizations need:

  • spatial joins between large datasets

  • automated enrichment of addresses, points, or polygons

  • region-level aggregation

  • geospatial quality checks

  • integration with data engineering workflows

  • reusable data products for BI, AI, or operational systems

4. Data platforms with geospatial capabilities

A modern data platform helps organizations manage spatial and non-spatial data together. Databricks fits this category when geospatial analytics needs to scale beyond isolated GIS workflows.

Databricks can be used to ingest, process, and govern geospatial datasets alongside business data. This matters because location questions usually require more than geometry. For example, service coverage analysis may need customer data, demand forecasts, logistics constraints, opening hours, capacity rules, and administrative boundaries.

A data platform is the right choice when you need:

  • large-scale batch or streaming processing

  • central governance and lineage

  • integration with enterprise data

  • repeatable pipelines

  • advanced analytics or machine learning

  • shared outputs for BI, GIS, APIs, and applications

In this setup, Databricks becomes the geospatial backbone. GIS tools can still be used for expert workflows, while Databricks handles scalable transformation, enrichment, and analytical production.

How to choose: GIS tool, data platform, or combination?

The right choice for your teams and company should be based on workload, users, and operating model. GIS tools and data platforms are the real either/or choice here; they are the foundation everything else sits on. Mapping and GeoApps, along with spatial databases and engines, aren’t a separate fork in that decision. They’re layers that plug into whichever foundation you pick, so it’s worth knowing where each one lands before working through the choice below.

Where mapping and GeoApps fit. This layer is about who needs to see the output, not which foundation produced it. A dashboard or web map works the same way whether the data behind it comes from a GIS system or from Databricks. Pick mapping software based on your audience and use case, not as an alternative to the foundation decision.

Where spatial databases and engines fit. This category rarely gets chosen on its own either. It’s usually already built into whichever foundation you pick: GIS systems often ship with something like PostGIS underneath, and Databricks includes its own spatial engine and functions. You inherit this layer from your foundation choice rather than selecting it separately.

With that settled, here’s how to decide on the foundation itself.

Choose a GIS tool when spatial experts are the main users

A GIS-first approach fits when the work is mainly visual, specialist, and layer-based. This includes editing boundaries, checking geometry quality, preparing maps, maintaining authoritative spatial datasets, and performing expert analysis.

The strength of GIS is human interpretation. If the main challenge is understanding the spatial context and manually validating layers, GIS software should stay central.

Choose a data platform when scale and automation matter

A data-platform approach fits when the work must run repeatedly, serve many consumers, or combine spatial data with broader enterprise data.

Choose Databricks when the use case includes:

  • high-volume GPS, IoT, or mobility data

  • daily or real-time pipeline refreshes

  • customer, asset, network, or operational data integration

  • governed access to sensitive location data

  • spatial features for machine learning

  • large point-in-polygon or grid-based analysis

  • consistent outputs for multiple downstream tools

This is where Databricks creates value as a scalable geospatial backbone, not by replacing every spatial tool, but by making geospatial processing production-ready.

Choose a combination when both expert workflows and scalable analytics are needed

Most mature organizations need both. GIS teams need specialist tools. Data teams need governed pipelines. Business teams need reliable outputs in dashboards, maps, and applications.

A combined architecture usually works best:

  • GIS tools manage specialist editing and spatial expertise

  • Databricks processes and governs large-scale geospatial data

  • Mapping and BI tools present curated outputs

  • APIs and data products make spatial insights reusable

  • Governance standards apply across the full data flow

This model avoids two common problems: GIS teams becoming responsible for enterprise-scale data engineering, and data teams underestimating spatial complexity.

Knowing which architecture fits is one thing. Getting there in the right order is another, and that’s where most projects actually go wrong.

A practical process for getting there

A strong geospatial solution starts with architecture, not tooling. Use the following process to shape the right landscape.

Step 1: Define the business question. Start with the decision the organization needs to improve. Examples include service coverage, risk exposure, route performance, asset planning, regional demand, or location allocation. Avoid starting with a map request. A map may be the output, but the value comes from the decision it supports.

Step 2: Identify spatial and non-spatial data. List the required datasets. This may include coordinates, addresses, polygons, routes, administrative boundaries, sensor records, satellite imagery, remote sensing feeds, customer data, operational systems, and external data. Check ownership, update frequency, quality, and sensitivity. Location data can be commercially or personally sensitive, so governance should be considered early.

Step 3: Decide where processing should happen. Not every spatial task belongs in the same tool. Editing and visual validation may stay in GIS. Heavy transformations, joins, and enrichment may move to Databricks. Visualization may happen in mapping software, BI tools, or applications. This separation gives each tool a clear role.

Step 4: Build reusable spatial data products. A spatial data product is a trusted dataset or metric that can be reused. Examples include service-area assignments, regional demand indicators, asset exposure scores, or H3-indexed mobility aggregates. Reusable data products reduce manual work and improve consistency.

Step 5: Operationalize and govern. Production geospatial pipelines need tests, monitoring, access control, documentation, and clear ownership. A successful proof of concept should not remain a one-off notebook or manual export, because nobody monitors, tests, or maintains a notebook the way they would a production system. Databricks supports this operating model by bringing geospatial analytics into the same platform used for data engineering, BI, AI, and governance.

Common mistakes in geospatial projects

Choosing the right foundation doesn’t guarantee a smooth rollout. These are the mistakes that show up even in organizations that got the GIS-versus-Databricks decision right, worth checking against before you start building.

Mistake 1: Treating GIS as the full data platform. GIS software is valuable, but it should not always be the central system for all analytical processing. When location data must be combined with enterprise datasets, a lakehouse architecture is often more sustainable.

Mistake 2: Removing GIS from the process. Data platforms do not replace spatial expertise. Coordinate systems, geometry validity, spatial relationships, and map interpretation require domain knowledge. GIS teams should be part of the design.

Mistake 3: Building isolated solutions. A one-off map, script, or dashboard may solve an immediate question, but it often creates a maintenance problem down the line: nobody remembers how it was built once its creator moves on, and it quietly breaks the next time a source system changes. Reusable data products and governed pipelines create longer-term value instead.

Mistake 4: Ignoring non-specialist users. Not every user is a GIS specialist. Good geospatial solutions translate spatial analysis into clear outputs: dashboards, alerts, APIs, reports, or simple map interfaces.

Mistake 5: Underestimating performance. Spatial joins and proximity calculations can become expensive at scale. Grid indexing, partitioning, spatial functions, and data-model choices should be part of the technical design from the start.

Knowing these mistakes in advance helps. Avoiding them under real deadline pressure is harder, and that’s usually where a partner with hands-on delivery experience makes the difference.

RevoData’s approach to geospatial solutions

As a trusted partner, RevoData helps organizations navigate complex geospatial issues by connecting GIS knowledge with modern data-platform engineering on Databricks.

As a Databricks Gold Partner, RevoData brings certified platform expertise to geospatial challenges. RevoData’s Databricks-certified consultants help organizations design scalable architectures, accelerate implementations, and apply quality standards to production pipelines. RevoData also invests strongly in continuous learning and has one of the highest numbers of Databricks Champions in EMEA.

The practical focus is clear: keep GIS where specialist spatial work belongs, use Databricks for scalable geospatial processing, and deliver trusted outputs to business users.

Need help choosing between GIS tooling, Databricks, or a combined geospatial architecture? Contact RevoData for consultancy around your geospatial data needs.

Final thought

The real decision here isn’t maps versus data platforms. It’s designing a landscape where geographic information becomes reliable, scalable, and useful across the organization.

A GIS tool is often the right choice for spatial experts. A data platform is the right choice for scale, governance, and integration. A combination is often the strongest answer.

With Databricks as the geospatial backbone, organizations can keep the strengths of GIS while making location intelligence part of their broader data and AI strategy.

FAQ's

RevoData helps organizations design and implement geospatial solutions on Databricks. This includes architecture, data engineering, spatial processing, governance, BI integration, and production-ready pipelines. The focus is on combining GIS expertise with scalable data-platform capabilities.

Use a GIS tool for specialist editing, map production, and expert spatial workflows. Use Databricks when geospatial data must be processed at scale, joined with enterprise data, governed centrally, or used in analytics and machine learning. Many organizations benefit from using both.

Common categories include GIS software, mapping tools, spatial databases, geospatial processing engines, data platforms, BI tools, and APIs. In a modern architecture, these tools should not operate as disconnected systems. They should share trusted data products and clear ownership.

Start with the decision the user needs to make. Limit the number of layers, use clear labels, define metrics carefully, and avoid unnecessary technical detail. The underlying data preparation can happen in Databricks, while the final map or dashboard presents only the information users need.

Databricks can act as the scalable geospatial backbone. It supports ingestion, transformation, enrichment, large-scale analysis, governance, and integration with BI or AI workflows. GIS tools can remain the expert interface, while Databricks handles repeatable processing.

Yes. Geospatial analytics increasingly overlaps with data engineering, cloud platforms, AI, and BI. RevoData works with consultants who combine technical depth with continuous learning. Candidates interested in data-platform architecture and Databricks can explore RevoData vacancies.

]]>
Geospatial Software Overview: Which Tools Fit a Modern Data Platform? https://revodata.nl/geospatial-software-overview-which-tools-fit-a-modern-data-platform/ Fri, 07 Aug 2026 07:57:17 +0000 https://revodata.nl/?p=7710

Every day, location data quietly makes or breaks a decision: a delivery route runs 20 minutes longer than scheduled, a new store opens in the wrong zip code, a flood risk gets flagged three weeks after the fact because someone had to manually cross-reference three spreadsheets and a shapefile. Geospatial analysis always mattered; what’s changing is the amount of data being collected and the variety of data types. Now, more teams than ever outside the GIS department use geospatial data for analysis and insights.

Location data now also shows up in customer analytics, logistics planning, site selection, mobility data, climate risk, telecom coverage, asset tracking, and public-sector reporting. This widening usage is exactly where most organizations get stuck. The GIS tools that help build their spatial expertise were never designed to process billions of GPS points, join polygons with enterprise data, or feed a machine learning model in production.

This article breaks down what geospatial software actually covers, where the categories differ, and how to combine classic GIS tools with Databricks instead of choosing between them for actionable intelligence. The short version: for most mature teams in the geospatial industry, the strongest setup is GIS tools for specialist workflows and Databricks as the scalable backbone that turns spatial data into geospatial intelligence, creating insight worth acting on.

Ready to see what that looks like for your organization? Explore RevoData’s geospatial services for Databricks.

Commercial solutions, open source, or Databricks: which one actually fits?

Before going deeper, let’s have a quick overview of the most popular options in the geospatial industry: commercial GIS solutions, open-source geospatial software, and Databricks.

Commercial tools (e.g. ArcGIS) Open source (e.g. QGIS, PostGIS) Databricks
Best for Enterprise-grade editing, cartography, and authoritative spatial data management Smaller teams, ad hoc analysis, and transactional spatial workloads under about 100GB Large-scale processing, repeatable pipelines, and joining spatial data with the rest of the business
Cost model License-based, scales with seats and server infrastructure Free to use, cost shows up in self-hosting and in-house expertise Consumption-based, scales with compute and data volume
Where it runs Desktop, enterprise server, or ArcGIS Online Desktop (QGIS) or a self-hosted database (PostGIS) Cloud lakehouse (AWS, Azure, GCP)
Scale ceiling Strong for standard enterprise data, slows down on billions of records or streaming data Solid up to tens of gigabytes, strains at genuine big-data volume Built for exactly that scale, distributed processing across a cluster
AI/ML readiness Improving through GeoAnalytics Engine and Esri's own AI models, but AI still sits outside the core GIS product Depends entirely on what you build on top yourself Native. Spatial data sits next to the ML tooling already, no separate environment needed

As you may have noticed, none of these three compete for the same job. They sit at different layers of the same stack; none of these replace each other. Commercial and open source GIS are where spatial expertise and authoritative data live. Databricks, on the other hand, is where that data goes to scale, gets governed, and feeds AI. The next sections cover exactly how those pieces connect.

What is geospatial software?

Geospatial software stores, processes, analyzes, and visualizes anything with a location component; examples of this are coordinates, addresses, routes, boundaries, grid cells, building footprints, parcels, service areas, or sensor positions.

In practice, geospatial tools are used for:

  • importing spatial formats such as GeoJSON, WKT, WKB, shapefiles, GeoPackage, or raster data

  • transforming coordinate reference systems

  • joining points, lines, and polygons

  • calculating distance, containment, overlap, buffers, and routes

  • producing maps and dashboards

  • enriching business data with location context

  • preparing spatial data for reporting, forecasting, or machine learning

The real value of these tools is not just in technical operations but also in the business insights they can unlock. These insights can help answer questions such as: Where should the next service location open? Which assets are exposed to flood risk this year, not five years ago?

A skilled GIS analyst can answer these questions with ease; the problem, however, is scale. A data platform needs to run automatically, daily, and across regions, not just when someone has time to open the desktop tool, to keep insights timely and relevant. This gap is where the software choice stops being a technical detail and starts being a strategic one.

The three types of geospatial software, and where each one runs out of road

Esri, open source GIS, and Databricks aren’t three unrelated options. They sit inside a broader landscape of three functional categories; knowing which category each one belongs to makes it easy to understand what you need.

Geographic information system (GIS) software (Esri and open source alternatives)

GIS software, and the broader systems built around it, from spatial databases to map services to publishing workflows, is where spatial expertise lives. This is the category to which both Esri and open source tools like QGIS belong. Whether it’s a proprietary platform like Esri’s ArcGIS or an open source option, this category is built for map creation, spatial editing, coordinate systems, topology checks, and the kind of visual, hands-on analysis that government, engineering, utilities, environmental, real estate, telecom, and transport teams rely on every day. If you need someone to inspect, correct, draw, classify, or make a judgment call on spatial data, GIS is still the best interface for that job.

Where you might run into trouble is that most GIS software and systems weren’t built for the current state of technology, nor for where technology is heading. These tools were not made for streaming data at volume, automated batch pipelines, lakehouse governance, or ML workflows. The moment spatial data needs to be integrated with ERP, CRM, IoT, finance, or predictive models, GIS stops being the whole answer and becomes one piece of a bigger, more robust architecture.

Mapping software

This category sits outside the commercial/open source/Databricks comparison above; it’s a distinct layer none of those three are primarily built for. Mapping software is built to communicate, not to process. Tools like Mapbox, the Google Maps API, and other mapping services turn spatial data into interactive maps, embedded map layers, and shareable dashboards. A good map does more work than a spreadsheet ever could: it makes clusters, outliers, service gaps, regional trends, and bottlenecks all instantly visible and easy to use for a broader group.

The problem with mapping software is that it’s not great of data driven desicion making as the data behind it does not get automatically validated or refreshed. In a modern setup, maintenance work happens in Databricks: data gets refreshed on a schedule instead of whenever someone remembers, quality checks catch bad or incomplete records before they ever reach the dashboard, and there’s a clear record of when the data last changed and where it came from. Mapping software remains the last visible step, not the place where the real answer gets built.

Spatial databases and geospatial engines (where Databricks fits)

This is the category built for repeatability and intelligence: spatial SQL, indexing, and joins that run automatically as part of a data pipeline, rather than requiring someone to manually complete these actions. A traditional spatial database handles this well up to a point, but it runs on a single server, and a single server can only index, search, and join so much data before it slows down and eventually can’t keep up at all. Databricks, and cloud data platforms like it, remove that ceiling by spreading the same work across a distributed cluster (multiple machines working together instead of one), so the system scales with the data instead of grinding to a halt the moment a dataset outgrows what one machine can handle.

Why classic GIS and Databricks need each other

The mistake a lot of organizations make is assuming that a modern data platform means replacing their classic GIS. Modernising means giving each tool the job it is best at to ensure you can meet technical business needs.

Databricks is the foundation that makes scaling and GeoAI possible. It’s great for ingesting high-volume batch and streaming data, storing spatial and business data in one lakehouse instead of two silos, running repeatable transformations, joining spatial data with everything else the business knows, and feeding BI, AI, and ML from a single governed foundation. This combination gives those working with the data the flexibility and creativity for thorough data processing, analysis, and visualisation. While end users get current and robust geospatial insights to make solid decisions based on.

GIS is where spatial expertise shines, the foundation that makes accuracy and trust possible. It’s great for editing geometries, validating layers visually, working with projections and cartography, publishing operational map layers, and supporting the people who genuinely think in space. This precision gives spatial specialists the control and confidence to get boundaries and details right the first time. While the teams who depend on that data get maps and records they can trust as the authoritative source of truth.

Split the responsibility this way, and everyone wins. GIS specialists keep the precision tools they need, data engineers get a backbone that can actually run in production, and business users get consistent answers. That combination is what turns raw spatial data into geospatial intelligence in the first place.

How to actually connect classic GIS to Databricks

Here’s what the integration looks like in practice, depending on which GIS stack you’re running.

If you’re on Esri

Esri and Databricks have a formal partnership, and there are three real ways to connect them, not just one:

  1. ArcGIS GeoAnalytics Engine. Esri’s own Spark plugin, installed directly on a Databricks cluster. It runs spatial SQL and analysis tools natively inside Databricks notebooks at Spark scale, and can write results straight back to ArcGIS Online or ArcGIS Enterprise as hosted feature layers. Use this when the heavy processing needs to happen in Databricks but the output still needs to live in ArcGIS for your GIS team.

  2. ArcGIS Data Pipelines. A low-code, visual ETL tool inside ArcGIS with a native Databricks connector, reading directly from Delta Lake tables without custom scripting. This is the simpler route when you mainly need ArcGIS kept in sync with data that already lives in your lakehouse.

  3. The ArcGIS API for Python inside Databricks notebooks. Lets your team query, manage, and pull ArcGIS Online or Enterprise content directly from a Databricks notebook. Useful when a data science workflow needs ArcGIS’s authoritative layers as an input, rather than the other way around.

A manual JDBC or ODBC connection between the two is technically possible but isn’t officially supported end to end. Most teams get more reliable results from one of the three patterns above.

If you’re on open source GIS

QGIS and PostGIS don’t have a formal partnership with Databricks, but the integration is arguably more direct:

  1. Native spatial SQL in Databricks. Recent Databricks runtimes support standard spatial functions (ST_Contains, ST_Distance, ST_Buffer, and the rest of the OGC set) directly in SQL, running on Databricks’ own vectorized query engine. The syntax closely mirrors PostGIS, so teams already comfortable with PostGIS have a short learning curve.

  2. Apache Sedona or Databricks’ own spatial libraries. For workloads native spatial SQL doesn’t cover, these add Spark-native spatial joins, H3 indexing, and large-scale geometry processing on the same cluster.

  3. A direct JDBC or ODBC bridge to PostGIS. For teams with an existing PostGIS database, Databricks can query it directly, keeping PostGIS as the system of record while letting Databricks join that data with everything else in the lakehouse.

QGIS itself remains a desktop tool in this picture. Data usually moves between QGIS and Databricks as files (GeoParquet, GeoJSON, shapefiles) rather than through a live connection, since QGIS is built for visual, single-user editing, not pipeline integration.

The pattern either way

Whichever GIS stack you’re on, the shape of the integration is the same. The GIS platform is the place where spatial experts edit and validate data. Databricks becomes the place where data gets processed at scale, joined with the rest of the business, governed centrally, and made available to AI. The tools change while the division of labour stays the same.

Geospatial AI: what Databricks makes possible that GIS alone can’t

This is where geospatial intelligence actually earns its name. A map can show you where something happened. AI on top of a governed spatial platform can tell you where something is likely to happen next.

GIS software was built for a person to look at a map and make a judgment call. It wasn’t built to train a model, score millions of records against that model, or refresh those scores automatically as new data arrives. Once spatial data sits in a governed lakehouse next to the rest of the business’s data, a few things become possible that simply aren’t practical in a GIS-only environment:

  • Predictive risk and demand models. Combine location with historical, weather, or operational data to forecast flood exposure, churn by branch, demand by micro-region, or maintenance needs before they become incidents, instead of reacting after the fact.

  • Computer vision on satellite, aerial, and LiDAR imagery. Detect land-use change, infrastructure damage, crop health, or unauthorized construction across thousands of images automatically, at a scale no team could review manually.

  • Anomaly detection on movement and sensor data. Flag unusual routes, unexpected dwell times, or sensor drift in near real time, rather than discovering the pattern weeks later in a quarterly report.

  • Spatial features feeding broader ML pipelines. Distance to nearest facility, service-area density, and proximity clustering can be engineered once as governed features and reused across multiple models, instead of recalculated by hand for every new project.

None of this replaces the GIS analyst’s judgment. Someone still needs to validate the boundaries, sanity-check the imagery labels, and decide what the model’s output actually means for the business. What changes is the volume of ground that judgment can now cover. A spatial team that used to answer one question at a time can now supervise a system that answers thousands, continuously.

What this actually looks like in practice

Location allocation and service coverage. Retailers, logistics platforms, healthcare networks, and public-service organizations constantly need to decide which locations serve which demand. GIS helps experts sanity-check the result visually. Databricks processes the underlying data at scale (customer locations, travel zones, capacity, delivery constraints, demographics, historical demand) to actually drive territory planning and branch optimization.

Mobility and GPS data. Millions or billions of GPS records will break a desktop workflow before lunch. Databricks cleans the raw data, removes duplicates, maps events to zones, aggregates trips, and calculates dwell times, while mapping software visualizes the output and GIS experts dig into the anomalies that actually need a human eye.

Risk, climate, and asset exposure. Combining asset locations with hazard zones, boundaries, and elevation data is only useful if you can refresh it, not just check it once. Databricks turns a one-time exposure check into a pipeline that updates the full portfolio automatically.

Telecom and network planning. GIS shows coverage and helps plan interventions. Databricks joins that network data with customer, usage, and operational data to actually prioritize where the next investment should go.

Public sector and urban analytics. GIS systems stay the trusted source for spatial maintenance and publication, while Databricks builds the governed analytical layer that ties spatial data into permits, mobility, infrastructure, and policy reporting.

Choosing the right geospatial software: 5 questions to ask first

1. Who’s actually using it? GIS analysts need editing and cartography. Data engineers need APIs, notebooks, and orchestration. Business users need a dashboard, not a shapefile. No single tool serves all three well. Design for each group instead of forcing one interface on everyone.

2. How much data, and how fast does it move? Standard GIS software handles small datasets fine. Nationwide address data, GPS streams, or recurring polygon joins need something built to scale. Move the heavy lifting to Databricks and send curated results back to GIS and BI.

3. How much does it need to talk to the rest of the business? The more location data depends on customer, finance, or logistics data, the more it needs to sit close to the enterprise platform. Otherwise you’re maintaining manual exports and inconsistent definitions forever.

4. What’s the governance and security risk? Addresses, movement data, and infrastructure locations are often sensitive data. Governance and security have to be built in from the start, not bolted on once something goes wrong.

5. How deep does the analysis need to go? If the roadmap includes AI, forecasting, or real-time processing, that’s your answer right there: a data platform approach, not just a better map.

Five mistakes that undo a geospatial modernization

  • Treating GIS as a data warehouse. It’s the source of truth for specific layers, not the home for enterprise-wide analytics.

  • Ripping out GIS too fast. Spatial data quality still depends on domain expertise. Keep GIS where it earns its place and move scale-heavy work to Databricks.

  • Skipping spatial indexing. Comparing every geometry to every other geometry gets expensive fast. Grid-based indexing keeps large-scale analysis practical.

  • Shipping notebooks, not pipelines. A notebook proves a method works. Production needs tests, monitoring, access control, and an owner.

  • Keeping spatial data in its own corner. The moment location joins the rest of the data platform, it becomes reusable across BI, AI, operations, and GIS, instead of a one-off asset nobody else can touch.

RevoData’s approach: Databricks as the geospatial backbone

RevoData provides GIS consulting and geospatial services that connect your existing GIS tools to scalable data engineering. RevoData helps you integrate, configure, and get more out of what you already have.

As a Databricks Gold Partner with 100% Databricks-certified consultants, RevoData brings the platform depth to turn spatial use cases into pipelines that actually run in production, not just in a demo. Most geospatial projects stall in the gap between departments: GIS teams know the spatial logic, data teams know scale and governance, and business teams just want a reliable answer. RevoData builds the architecture that closes that gap.

Planning to scale geospatial analytics on Databricks? Talk to RevoData about connecting your GIS workflows, lakehouse architecture, and business outcomes.

The recommended architecture, in one list

  • GIS software for specialist editing, map production, and domain workflows

  • Databricks for ingestion, transformation, spatial joins, H3 indexing, data quality, governance, and ML

  • Mapping software or BI tools for consumption and communication

  • APIs or data products for operational applications

  • Clear ownership split between GIS, data platform, and business teams

Get this right and GIS stays exactly as valuable as it always was. It just stops being a bottleneck for everything downstream.

The bottom line

Modern geospatial software isn’t a single product decision. It’s an architecture decision.

Classic GIS still earns its place for spatial experts. Mapping software still earns its place for communication. Databricks adds what neither can provide on its own: a scalable backbone for production pipelines, governed data products, and geospatial AI, built to keep up with how much location data your organization generates now.

Ready to stop choosing between GIS and scale? Explore RevoData’s geospatial services and see how Databricks becomes the backbone for your spatial analytics.

FAQ's

Software for working with anything that has a location component: coordinates, addresses, boundaries, routes, and areas. Core functions include mapping, spatial joins, distance calculations, containment checks, routing, enrichment, and spatial data quality control.

Match the tool to the workload. Choose GIS software for editing and expert spatial analysis, mapping software for visual communication, and Databricks (or another scalable data platform) once you’re dealing with large datasets, repeatable pipelines, enterprise integration, governance, or AI. Most organizations end up needing more than one.

They make location measurable and visible, supporting decisions in logistics, infrastructure, planning, environment, retail, telecom, and risk. Connected to a data platform, GIS outputs also become reusable inputs for dashboards, models, and operational processes instead of one-off exports.

“Best” depends entirely on your use case, existing architecture, licensing, data volume, and users. The more useful question isn’t which desktop GIS product wins. It’s how well your GIS workflows connect to scalable, governed analytics.

For high-volume transformation, enrichment, spatial joins, and analytics pipelines, yes, often. For expert spatial editing, cartography, and visual inspection, no, and it shouldn’t try to. The strongest setup uses Databricks as the backbone and GIS as the specialist interface.

Spatial data rarely stays spatial-only for long. It needs to be processed alongside enterprise data, governed centrally, and reused across BI, AI, and operations. Databricks supports the scalable processing, spatial functions, and H3-based indexing needed to turn one-off spatial analysis into a repeatable data product.

Geospatial AI applies machine learning to location data: predicting risk, detecting patterns in movement or sensor data, or analyzing satellite and aerial imagery at scale. It depends on having spatial data in a platform built for AI workloads in the first place, which is exactly the gap Databricks fills alongside traditional GIS tools.

]]>
Databricks Academy training: How to become a Databricks Expert https://revodata.nl/databricks-academy-training-how-to-become-a-databricks-expert/ Wed, 05 Aug 2026 14:18:21 +0000 https://revodata.nl/?p=7706

Learning Databricks is easier when the training path matches your role. A data engineer does not need the same first steps as a BI analyst. A data scientist does not need to start with the same certification as a platform architect. Databricks Academy helps professionals build structured knowledge, but real expertise comes from applying that knowledge in practical projects: pipelines, dashboards, machine learning workflows, governance, and production operations.

For organizations, the question is not only “Which Databricks course should our team follow?” The stronger question is: “Which skills do we need to run Databricks well in our own environment?” That is an important distinction. A certificate can validate knowledge, while practical training makes teams confident in real work.

Databricks Academy is Databricks’ own self-paced training environment, great for structured learning at your own speed. RevoData also runs its own small-group, instructor-led Databricks Training sessions directly, built around your own data landscape, alongside certification preparation and hands-on enablement from 100% Databricks-certified consultants. For teams that need extra capacity, Managed Databricks can also act as an extension of the internal team.

Want a practical Databricks learning path for your team? RevoData can help you combine self-paced Databricks Academy content, certification preparation, and RevoData’s own instructor-led Databricks Training.

What is Databricks Academy?

Databricks Academy is the official learning environment for Databricks training. It offers courses, learning paths, and certification preparation for professionals working with data engineering, analytics, machine learning, AI, and platform administration.

Databricks courses are useful because the platform covers several disciplines. A modern Databricks environment may include data pipelines, SQL analytics, machine learning, generative AI, governance, orchestration, notebooks, dashboards, and cloud integration. Without a structured learning path, teams often learn fragments of the platform without understanding how the pieces work together.

Databricks Academy helps learners build a foundation in areas such as:

  • the Databricks workspace

  • notebooks and SQL

  • data ingestion and transformation

  • Lakehouse architecture

  • Delta Lake concepts

  • data engineering workflows

  • BI and dashboarding

  • machine learning and AI workflows

  • governance and access control

  • certification preparation

The Academy is a strong starting point, but it should not be the only learning method. The most valuable skills are built by solving realistic tasks: loading data, cleaning it, modeling it, testing pipelines, managing permissions, controlling cost, and delivering outputs to users.

Which learning path fits your role?

Different roles need different Databricks skills. A good training plan should separate common foundations from role-specific depth.

Learning path for data engineers

Data engineers need to build reliable data products. Their learning path should focus on pipelines, orchestration, data quality, performance, and maintainability.

A practical Databricks training path for data engineers should cover:

  • Databricks workspace basics

  • Python and SQL in notebooks

  • PySpark fundamentals

  • ingestion patterns

  • Auto Loader and pipeline design

  • Delta Lake concepts

  • Lakeflow and orchestration

  • medallion architecture patterns

  • data quality checks

  • job scheduling and monitoring

  • Git-based development workflows

  • cost-aware compute usage

  • governance with catalogs, schemas, and permissions

For data engineers, the most relevant certification route often starts with a Databricks data engineer certification at the associate level. More experienced engineers can then move toward advanced or professional-level validation when they have enough practical platform experience.

A separate Python course can also be useful before or alongside Databricks training. Python is not the only language used on Databricks, but it is widely used for data engineering, notebooks, PySpark, and automation.

Learning path for data analysts and BI specialists

Analysts need to turn trusted data into clear insights. They usually do not need the same depth in distributed processing as engineers, but they do need strong SQL, data modeling awareness, and confidence with governed datasets.

A Databricks learning path for analysts should cover:

  • workspace navigation

  • SQL querying

  • dashboards and visual analysis

  • working with curated datasets

  • understanding Lakehouse concepts

  • basic data quality interpretation

  • collaboration with data engineers

  • permissions and governance basics

  • performance-aware querying

  • metric definitions and data product usage

For BI specialists, the key skill is knowing how Databricks fits into the analytics chain. Databricks may prepare and serve the data, while BI tools present it to business users. Analysts should understand where the data comes from, which transformations were applied, and which definitions are trusted. The Databricks Certified Data Analyst Associate route can be relevant for analysts who want to validate their platform knowledge.

Learning path for data scientists

Data scientists need to move from experimentation to production-quality machine learning. Databricks is useful because it connects notebooks, data preparation, model development, tracking, deployment, and monitoring.

A data science training path should cover:

  • Python on Databricks

  • working with notebooks

  • exploratory data analysis

  • feature preparation

  • MLflow concepts

  • experiment tracking

  • model training and evaluation

  • AutoML where appropriate

  • model deployment patterns

  • governance for data and models

  • collaboration with data engineering teams

  • responsible use of AI and machine learning

For data scientists, certification can be useful, but practical model lifecycle skills matter more than exam preparation alone. Training a model is the easy part. The harder, more valuable skill is building workflows that can be tested, governed, and maintained.

Learning path for AI engineers and app developers

AI engineers and app developers need to understand how Databricks supports AI applications, agents, and workflows connected to enterprise data.

A practical AI learning path should cover:

  • data preparation for AI applications

  • vector search and retrieval patterns

  • foundation model usage

  • evaluation of generated outputs

  • prompt and tool design

  • governance and access control

  • monitoring

  • integration with applications

  • security and data privacy considerations

For app developers specifically, Databricks knowledge becomes more relevant when applications depend on trusted enterprise data or AI outputs. The developer does not need to become a full data engineer, but should understand how to consume governed data products and AI services safely.

Learning path for platform owners and architects

Platform owners need to run Databricks responsibly. Their learning path should focus on architecture, governance, security, cost management, and operating models.

Important topics include:

  • workspace strategy

  • identity and access management

  • Unity Catalog concepts

  • compute policies

  • cost controls

  • environment separation

  • deployment standards

  • data governance

  • monitoring

  • platform support

  • adoption planning

This role is often underestimated. A team can complete several Databricks courses and still struggle if platform ownership is unclear. RevoData often sees that adoption improves when technical enablement is paired with clear standards and support.

Learning path for data stewards

Data stewards are responsible for the trustworthiness of the data itself: quality, compliance, and consistent definitions across the organization. Their learning path should focus on governance concepts more than pipeline engineering.

A Databricks learning path for data stewards should cover:

  • Unity Catalog fundamentals

  • data classification and sensitivity labeling

  • access control and permission models

  • data quality rules and monitoring

  • lineage and audit trails

  • compliance requirements relevant to the organization

  • collaboration with data engineers and platform owners on governance standards

Data stewardship often gets treated as a side responsibility rather than a distinct skill set. That’s a mistake: without someone actively owning data quality and compliance, governance rules exist on paper but don’t get enforced in practice.

Learning path for geospatial engineers

Geospatial engineers bring spatial expertise into the same platform used for the rest of the organization’s data. Their learning path should combine GIS fundamentals with Databricks-native spatial tools.

A Databricks learning path for geospatial engineers should cover:

  • GIS fundamentals and spatial data types

  • spatial functions and geospatial engines such as Apache Sedona

  • H3 indexing and spatial joins at scale

  • integrating existing GIS tools with Databricks pipelines

  • remote sensing and satellite imagery workflows

  • photogrammetry and 3D information extraction

  • governance for spatial and location-sensitive data

This is a newer, more specialized track, and one RevoData offers and has particular depth in, given its work connecting classic GIS tools to Databricks as a scalable geospatial backbone.

Knowing which path fits a role is only useful once someone actually walks it. Here’s how to put that into practice.

How to start with Databricks Academy

A practical first step is to avoid starting with the exam. Start with the role and the work.

Step 1: Define the role. Choose the learning path based on the learner’s actual responsibilities. Is the person building pipelines, creating dashboards, training models, managing the platform, or building AI applications?

Step 2: Build a shared foundation. Before specializing, teams should understand the basics: what Databricks is, how the workspace works, how data is organized, and how collaboration happens.

Step 3: Use Databricks Free Edition for practice. Databricks Free Edition can be useful for learning and experimentation. It gives learners a no-cost environment to explore data and AI concepts. It is suitable for personal learning, prototyping, and experimentation, but it is not the same as an enterprise implementation with full governance, representative data, and production controls.

Step 4: Follow role-specific Databricks courses. After the foundation, choose courses that match the role. Engineers should go deeper into data pipelines. Analysts should focus on SQL and dashboards. Data scientists should focus on machine learning workflows. Platform owners should focus on governance and administration.

Step 5: Add hands-on assignments. Training becomes more useful when learners apply concepts immediately. Examples of assignments include building an ETL pipeline, creating a governed dataset, writing SQL queries, tracking an ML experiment, or publishing a dashboard.

Step 6: Prepare for certification. Certification preparation should come after practice. Learners who have only watched course material may recognize terms but struggle with applied questions. Hands-on use makes certification preparation more effective.

Step 7: Connect learning to team standards. Training should result in shared ways of working. Examples include naming standards, data quality expectations, Git usage, job scheduling patterns, cost controls, and governance rules.

That seven-step process holds regardless of which cloud a team runs on, but the cloud does change some of the details worth training for.

Training Azure Databricks: what changes?

Azure Databricks is Databricks integrated with Azure. For organizations already working on Azure, training should include both Databricks concepts and Azure-specific operating patterns.

Azure Databricks learning should cover:

  • workspace deployment and access

  • identity integration

  • storage patterns

  • networking and security

  • connection to Azure data services

  • cost and compute management

  • governance across cloud and Databricks layers

The platform skills remain Databricks skills, but the operating context matters. A learner who can build a notebook may still need support understanding enterprise security, networking, identity, and deployment. That is where a partner-led learning path can help. RevoData connects platform theory to the way an organization actually runs Azure and Databricks.

Once training accounts for the platform, the role, and the cloud it runs on, the last piece is proving that knowledge formally.

Databricks certifications: which one should you choose?

Databricks certifications validate role-specific knowledge. The right certification depends on the learner’s role and experience.

Typical routes include:

Data Engineer certification. Best for professionals who build and maintain data pipelines, transform datasets, manage reliability, and prepare data products for analytics or AI.

Data Analyst certification. Best for analysts and BI specialists who use SQL, dashboards, and governed datasets to produce insights.

Machine Learning certification. Best for data scientists and machine learning engineers who build, evaluate, and manage models on Databricks.

Generative AI Engineer certification. Best for AI engineers and app developers who design, build, and deploy generative AI solutions on Databricks, including retrieval-augmented generation, AI assistants, and agents. This maps directly to the AI engineer learning path above and is one of Databricks’ fastest-growing certification tracks.

Solution Architect or platform-oriented learning. Best for architects, platform owners, and senior consultants who design environments, governance models, and end-to-end solutions.

Certification is useful, but it should not become the only target. A certified professional should also be able to explain trade-offs, debug workflows, work with real data, and collaborate across teams.

Certification tells you what someone should know. It doesn’t guarantee the learning path that got them there avoided the usual pitfalls.

Common mistakes when learning Databricks

Starting with too much theory. Concepts matter, but Databricks is best learned by doing. Learners should build pipelines, run notebooks, query data, test workflows, and inspect errors.

Choosing the wrong certification. A data analyst does not need to start with an engineering-heavy path. A data engineer should not focus only on dashboarding. Match the certification to the work.

Ignoring Python and SQL fundamentals. Databricks training is easier when learners already understand SQL and basic Python. A Python course or SQL refresher can reduce friction.

Treating Free Edition as production training. Databricks Free Edition is useful for learning, but enterprise projects require additional knowledge: governance, access control, networking, cost management, and deployment standards.

Training individuals without enabling the team. One trained person can help, but Databricks adoption needs shared standards. Teams should agree on development patterns, review practices, quality checks, and support responsibilities.

Forgetting managed support. Not every organization needs to build every skill internally from day one. Managed Databricks can support platform reliability, governance, and best practices while internal teams build confidence.

That last point is worth its own explanation, since training and managed support solve different problems rather than one replacing the other.

RevoData’s Databricks Training service

Take your team beyond self-paced Databricks Academy content and general enablement with RevoData’s instructor-led Databricks Training offering. RevoData hosts small group in-house sessions with a live instructor for your company customized to your needs.

Example of trainings include:

  • Basic training– for teams new to Databricks who need fundamental platform knowledge

  • Data Engineer training -covering ingestion, transformation, storage, and ETL best practices

  • Machine Learning Engineer training– covering model development, evaluation, and deployment

  • Data Analyst training– covering advanced analysis and visualization techniques

  • Platform Engineer training– covering architecture, configuration, security, and resource optimization

  • Data Steward training– covering governance, compliance, and data quality management

  • Geospatial Engineer training– covering GIS fundamentals and hands-on geospatial tools in Databricks

This tends to matter most at three points: right after a new Databricks implementation, when someone changes roles or gets promoted into new platform responsibilities, and as a periodic refresher to keep a team current as Databricks itself keeps changing.

Self-paced Databricks Academy content is good for learning concepts at your own pace, and it’s worth using regardless of who delivers your team’s training. What an instructor-led session with RevoData adds on top of that is context: as a Databricks Gold Partner with 100% Databricks-certified consultants, RevoData’s trainers bring real implementation experience into the room, answer questions specific to your own data landscape on the spot, and adjust pace and depth as the session goes.

RevoData’s hands-on approach

RevoData helps teams learn Databricks through practical enablement, not theory alone. RevoData combines platform expertise with delivery experience.

The approach is role-based and hands-on:

  • Teams new to the platform learn through basic, foundational training before specializing

  • Data engineers learn by building pipelines

  • Analysts learn by working with governed datasets and SQL

  • Data scientists and machine learning engineers learn by developing reproducible ML workflows

  • Platform engineers learn by managing governance, cost, and standards

  • AI engineers learn by connecting data products to AI applications

  • Data stewards learn by setting up and enforcing real governance rules

  • Geospatial engineers learn by connecting GIS tools to Databricks pipelines

RevoData also supports organizations through Managed Databricks. This service can act as an extension of the internal team, helping with platform operations, best practices, troubleshooting, governance, and continuous improvement. That matters because training alone does not guarantee adoption. Teams need support while they apply new skills to real environments.

Want to combine self-paced Databricks Academy content with instructor-led enablement? RevoData can help design role-based training, certification preparation, and Managed Databricks support.

FAQ's

You can access Databricks Academy through the official Databricks training environment. Learners can create an account or use access connected to their organization. From there, they can browse available courses, learning paths, and certification preparation material.

Databricks offers role-based certifications for areas such as data engineering, data analysis, machine learning, and generative AI engineering. The right certification depends on your role. Engineers typically start with a data engineering path, analysts with a data analyst path, data scientists with a machine learning path, and AI engineers with the generative AI engineer path.

Databricks offers free on-demand training options for customers and learners, while certification exams and instructor-led training may have separate pricing. Costs can change, so always check the official Databricks training and certification pages before planning a budget.

Yes. Databricks Free Edition is useful for personal learning, experimentation, and prototyping. It is a good way to practice notebooks, datasets, AI, and machine learning concepts. For enterprise readiness, teams still need to learn governance, security, deployment, and cost management.

Python is highly useful, especially for data engineers, data scientists, and AI engineers. Analysts may start with SQL first. A Python course can help learners become more confident with notebooks, PySpark, and automation tasks.

Yes. RevoData delivers its own small-group, instructor-led Databricks Training sessions directly, alongside certification preparation and practical enablement built into consultancy engagements. Managed Databricks can also support teams that need expert help while building internal capability.

Yes. RevoData is a Databricks Gold Partner with 100% Databricks-certified consultants. The team has a strong focus on quality, continuous learning, and practical implementation support.

]]>
How to Start with Business Process Automation https://revodata.nl/how-to-start-with-business-process-automation/ Wed, 05 Aug 2026 13:24:50 +0000 https://revodata.nl/?p=7685

Finance teams spend hours checking invoices against purchase orders, customer service agents copy details from emails into a CRM system, and operations teams export data from one platform, clean it in spreadsheets, and upload it somewhere else. None of these tasks is difficult in isolation, but together they slow down the organization, introduce errors, and make reporting less reliable.

This is where business process automation comes in as a starting point. Business process automation is the structured use of software, data, and rules to reduce manual work in repeatable processes. The goal is not to remove people from the organization but to let people spend less time on repetitive handovers and more time on decisions, exceptions, and improvement.

Once that foundation is in place, the natural next step is data-driven AI automation. Traditional automation and robotic process automation (RPA), software bots that mimic repetitive human actions like copying data or clicking through screens, can only execute fixed steps. AI-driven automation goes further: it can interpret documents, classify requests, detect anomalies, recommend actions, and adapt to more complex inputs. Databricks can act as the central engine for these workflows by combining data engineering, AI, governance, and orchestration in one platform.

Want to identify which business processes are ready for automation? RevoData’s AI Engineering team can help you assess use cases, estimate value, and build a focused proof of concept.

What is business process automation?

Business process automation means using technology to execute, support, or improve recurring business processes. These processes may involve approvals, data entry, document checks, notifications, reporting, handovers, customer communication, or operational decisions.

A process is a good candidate for automation when it has a clear trigger, repeatable steps, defined inputs, and measurable outputs. For example:

  • A new invoice arrives

  • A customer submits a request

  • A sensor produces an alert

  • A sales opportunity reaches a certain stage

  • A report needs to be refreshed

  • A file is uploaded

  • A transaction needs validation

Automation can then perform one or more steps, such as extracting data, checking rules, routing the item, updating a system, creating a task, sending a notification, or preparing a recommendation.

There are different levels of business process automation:

  • Rule-based automation. The system follows fixed instructions. Example: “If the invoice amount is below €5,000 and the supplier is approved, route it to finance.”

  • RPA. Robotic process automation uses software bots to mimic human actions in user interfaces, such as copying data between systems or filling out forms.

  • Workflow automation. The system coordinates tasks across people and systems. For example: approvals, status changes, reminders, and escalations.

  • Data-driven automation. The system uses data pipelines, analytics, and business logic to make processes more reliable and measurable.

  • AI-driven automation. The system uses AI models to interpret information, classify cases, detect patterns, make predictions, or recommend next steps.

The most mature organizations combine these layers. They use simple automation where rules are enough, and AI where the process requires interpretation or prediction.

Why automate business processes?

Automating business processes is important because many organizations still depend on manual coordination between systems, which creates delays, errors, and hidden costs.

The benefits of automating business processes are usually visible in five areas.

1. Less manual work. Teams spend less time copying data, checking standard conditions, routing cases, or preparing recurring reports. This frees capacity for work that requires judgment.

2. Fewer errors. Manual retyping, spreadsheet exports, and repeated handovers increase the chance of mistakes. Automation helps standardize execution and reduce avoidable errors.

3. Faster cycle times. Automated workflows can run immediately when a trigger occurs. This shortens response times in areas such as finance, operations, customer service, and IT.

4. Better visibility. Automated processes create data about the process itself. Teams can measure backlog, lead time, exception rates, approval delays, and quality issues.

5. Stronger AI readiness. AI automation needs reliable data and clear workflows. By automating business processes on a governed data platform, organizations create the foundation for more advanced AI use cases.

Research supports the productivity potential, but it also shows that implementation quality matters. McKinsey estimated that generative AI, combined with other automation technologies, could add 0.5 to 3.4 percentage points annually to productivity growth. That value does not appear automatically; it depends on redesigning work, integrating data, and managing risk well.

From RPA to data-driven AI automation

RPA is often the first step in automating business processes. It is useful when a task is repetitive, structured, and stable. A bot can log in, copy information, click through screens, and update records.

RPA has limitations. If an interface changes, a field moves, a document format differs, or an exception occurs, the bot may fail. It also tends to work around system fragmentation rather than solving it. AI automation goes further. It can work with unstructured data and more complex decisions. Examples include:

  • reading and classifying customer emails

  • extracting fields from invoices or contracts

  • detecting unusual transactions

  • predicting which cases need urgent attention

  • summarizing documents for review

  • recommending the next best action

  • routing cases based on content, risk, and history

This does not mean every process should become fully autonomous. The strongest AI automation designs include human review at the right moments. Low-risk cases can be processed automatically. Medium-risk cases can receive an AI recommendation. High-risk cases should be escalated to a specialist.

Databricks supports this shift by providing a central environment for data pipelines, AI models, governance, monitoring, and workflow orchestration. Instead of building isolated bots around disconnected systems, organizations can build data-driven workflows on a governed platform.

Which processes are suitable for automation?

A process is suitable for automation when it meets several criteria. These criteria are:

  • High volume. Processes that happen many times per week or per day usually offer a better return. Examples include invoice checks, customer request routing, report preparation, ticket classification, and transaction monitoring.

  • Repetitive steps. The more repeatable the steps, the easier it is to automate. Even when some cases require human judgment, the standard parts can often be automated.

  • Clear business rules. If the process has known decision rules, thresholds, or routing logic, automation can execute those rules consistently.

  • Structured or semi-structured data. Processes based on forms, tables, documents, emails, or system events can often be automated. AI is especially useful when the input is semi-structured or unstructured.

  • Frequent errors or delays. Processes with many manual corrections, bottlenecks, or missed handovers are strong candidates. Automation can standardize the flow and expose where exceptions occur.

  • Measurable value. Good automation candidates have clear success metrics. Examples include hours saved, shorter lead times, fewer errors, improved response times, lower cost per case, or better data quality.

Practical examples of business process automation

Those criteria are easier to apply with concrete cases in front of you. Here’s what they look like across a few common business functions.

  • Finance: invoice and purchase-order checks. Automation can extract invoice data, compare it with purchase orders, check supplier rules, identify mismatches, and route exceptions. AI can help interpret invoice descriptions or detect unusual values.

  • Customer service: request classification. AI can classify incoming messages, detect urgency, summarize context, and route the request to the right team. Workflow automation can create tasks, update statuses, and send notifications.

  • Operations: incident prioritization. Operational alerts can be enriched with historical data, asset information, and risk scores. Automation can prioritize incidents and recommend action.

  • HR: onboarding workflows. Automation can coordinate account creation, document collection, training tasks, approvals, and status updates. AI can help summarize documents or answer employee questions from approved knowledge sources.

  • Sales and marketing: lead enrichment and routing. Automation can enrich leads, score them, assign them to teams, and trigger follow-up actions. AI can support segmentation, content personalization, or next-best-action recommendations.

  • Data operations: recurring reporting. Instead of manually preparing spreadsheets, Databricks can refresh datasets, validate quality, run transformations, and feed dashboards or downstream applications.

Step-by-step: how to start automating business processes

Seeing those patterns is one thing. Turning one of them into a working process on your own is another. The path there isn’t complicated, but it does need to happen in order, starting with figuring out which process to tackle first.

Step 1: List candidate processes. Start by interviewing teams and mapping repeated work. Look for manual copying, spreadsheet work, recurring checks, approval chains, status updates, and exception handling.

Ask practical questions:

  • Which tasks are repeated every day?

  • Where do delays occur?

  • Which steps depend on manual data entry?

  • Which reports take too long to prepare?

  • Where do errors regularly appear?

  • Which decisions use the same information repeatedly?

Step 2: Score processes by value and feasibility. Not every process should be automated first. Score each candidate on value and feasibility.

  • Value criteria: time spent, cost impact, customer impact, error reduction, compliance value, and scalability.

  • Feasibility criteria: data availability, process clarity, system access, rule stability, exception rate, and governance risk.

Start with a process that has meaningful value and manageable complexity.

Step 3: Decide the right automation type. Choose the right technology pattern for the right task and situation.

  • Use simple workflow automation for approvals and notifications

  • Use RPA for stable interface-based tasks

  • Use data pipelines when the process depends on data movement and transformation

  • Use AI when interpretation, classification, or prediction is needed

  • Use a combined approach when execution and intelligence are both required

Step 4: Prepare the data. Automation quality depends on data quality. Check whether the required data is complete, current, accessible, and governed. For AI automation, data preparation is even more important. Models need reliable inputs, and business users need confidence in the outputs.

Step 5: Build a proof of concept. A proof of concept should test a real process with representative data. Keep the scope narrow, but make sure the test includes enough complexity to be meaningful.

A good PoC should answer:

  • Does the automation work technically?

  • How much manual work can be reduced?

  • Which exceptions remain?

  • What are the risks?

  • Which data quality issues appear?

  • What is needed for production?

Step 6: Add governance and monitoring. Production automation requires controls. Define access rights, logging, quality checks, escalation rules, cost monitoring, and ownership. For AI workflows, also monitor model performance, output quality, drift, bias risks, and human override rates.

Step 7: Scale with a roadmap. After the first successful process, create a roadmap. Reuse components where possible: data pipelines, validation logic, AI models, governance patterns, and workflow templates.

That roadmap only holds up if each process runs on the right technology to begin with. Here’s how those categories actually break down

Tools and software for automation

There is no single best tool for every business process. The right landscape depends on the process. These are the tools worth considering in your journey:

  • Workflow tools. Useful for approvals, task routing, notifications, and status tracking.

  • RPA platforms. Useful for repetitive, interface-based tasks in legacy systems.

  • Low-code automation platforms. Useful for fast workflow development where processes are relatively simple.

  • Data platforms. Useful when automation depends on data integration, transformation, quality, and analytics.

  • AI and machine learning platforms. Useful when processes require prediction, classification, summarization, anomaly detection, or recommendations.

  • Databricks. Relevant when organizations want a central engine for data-driven and AI-driven automation. Databricks can support data ingestion, transformation, machine learning, generative AI, governance, and workflow orchestration. It is especially useful when automation must connect multiple data sources and scale beyond isolated bots.

Common mistakes in business process automation

Even a well-designed automation strategy runs into the same handful of problems in practice. These are the common ones worth checking before committing to a specific process:

  • Automating before understanding the process. If a process is unclear, automation can make the confusion bigger. Map the process first. Identify ownership, inputs, outputs, and exceptions.

  • Choosing technology before use cases. A tool-first approach often creates isolated pilots. Start with business value and process fit, then choose the technology.

  • Ignoring data quality. Automation fails when data is incomplete, inconsistent, or inaccessible. Data readiness should be assessed before building.

  • Using RPA where data integration is needed. RPA can be useful, but it is not always the best long-term architecture. If the process depends heavily on data, a platform-based approach may be more sustainable.

  • Removing humans too quickly. Some decisions need review. AI automation should include clear escalation rules and human oversight for high-impact cases.

  • Measuring only hours saved. Time savings matter, but automation also creates value through fewer errors, faster response, better compliance, improved data quality, and stronger process insight.

Knowing these pitfalls in advance helps. Avoiding them under real deadline pressure is harder, and that’s usually where a partner with hands-on delivery experience makes the difference.

RevoData’s approach: helping you identify and implement AI use-cases

RevoData helps organizations move from manual work and rule-based automation to data-driven AI automation. The starting point is use-case identification: which processes are suitable, what value can they create, and what architecture is needed?

As a Databricks Gold Partner with 100% Databricks-certified consultants, RevoData brings platform expertise and practical delivery experience. The team helps organizations assess automation opportunities, prepare data, design Databricks-based workflows, and build proofs of concept that seamlessly move toward production.

RevoData’s AI Engineering approach focuses on:

  • identifying high-value automation candidates

  • assessing data readiness

  • selecting the right mix of workflow automation, RPA, data pipelines, and AI

  • designing Databricks as the central engine

  • building scalable proofs of concept

  • adding governance, monitoring, and quality controls

  • supporting teams through implementation

This approach helps organizations avoid disconnected automation projects. Instead, business process automation becomes part of a broader data and AI capability.

Ready to automate business processes with a data-driven approach? Explore RevoData’s process optimization expertise and start with a focused use-case assessment.

FAQ's

Business process automation is the use of technology to execute or support recurring business processes. It can include workflow automation, RPA, data pipelines, and AI-driven automation. Examples include invoice routing, customer request classification, report generation, and operational alerts.

Companies automate business processes to reduce manual work, shorten cycle times, reduce errors, improve visibility, and make processes easier to scale. Automation also creates a stronger foundation for AI because workflows become more structured and measurable.

Suitable processes usually have high volume, repetitive steps, clear rules, available data, and measurable value. Examples include finance checks, customer service routing, reporting, HR onboarding, lead processing, incident triage, and data quality monitoring.

Start by listing candidate processes, scoring them by value and feasibility, selecting one focused use case, checking data readiness, and building a proof of concept. After that, add governance, monitoring, and a roadmap for scaling.

RPA automates fixed, repetitive tasks by following predefined steps, often through a user interface. AI automation uses models to interpret information, classify cases, detect anomalies, predict outcomes, or recommend actions. RPA is useful for stable tasks; AI automation is better for variable or data-rich processes.

The best tool depends on the process. Workflow tools are useful for approvals and task routing. RPA is useful for stable interface-based work. Data platforms are useful for data-heavy processes. AI platforms are useful when interpretation or prediction is needed. Databricks is relevant when automation depends on governed data, AI models, and scalable workflows.

]]>
AI Automation: From Rule-Based Tasks to Learning Workflows https://revodata.nl/ai-automation-from-rule-based-tasks-to-learning-workflows/ Wed, 05 Aug 2026 12:00:34 +0000 https://revodata.nl/?p=7654

Many companies have already automated parts of their operations. Invoices are routed automatically, forms are copied from one system to another, customer emails are classified by keyword, and reports are generated on a schedule. That type of automation can reduce manual work, but it often breaks when situations change. A new document layout, an unclear customer request, an unexpected data value, or a missing field can stop the process. Traditional automation is good at following rules; it is less effective when the work requires interpretation.

AI automation changes that. Instead of solely executing fixed instructions, AI-driven automation can interpret text, learn from data, classify situations, detect patterns, make recommendations, and support decisions. This moves automation from rule-based execution to adaptive workflows. For organizations, the opportunity does not lie in doing the same tasks faster. The bigger value lies in redesigning processes around data, AI models, and human review. Databricks acts as the engine for those types of complex workflows by combining data engineering, machine learning, generative AI, governance, and orchestration in one platform.

Need help identifying which processes are ready for AI automation? RevoData can help you assess opportunities, define a proof of concept, and build a scalable Databricks architecture.

What is AI automation?

AI automation is the use of artificial intelligence to automate tasks, decisions, or workflows that require more than fixed business rules. It combines automation with AI capabilities such as natural language processing, machine learning, computer vision, prediction, classification, anomaly detection, and generative AI. A simple automation follows an instruction such as: “If the invoice amount is below €5,000 and the supplier is approved, send it to finance.” AI-driven automation can handle more complex situations, such as: “Read the invoice, extract the relevant fields, compare it with purchase order data, identify unusual values, check whether the description matches previous purchases, and recommend whether a human should review it.” The difference is interpretation. Traditional automation works best when the process is stable and predictable, whereas AI automation is useful when the process involves unstructured information, variation, context, or probability. AI automation can support many business functions, including finance, operations, logistics, customer service, marketing, compliance, HR, data management, and IT operations.

What is the difference between RPA and AI automation?

Traditional automation is often referred to as robotic process automation, or RPA. RPA uses software bots to perform repetitive digital tasks. These bots can click buttons, copy data, fill forms, move files, and execute predefined workflows. RPA is useful when the task is repetitive, rule-based, and stable. Examples include copying data from emails into a system, downloading reports, updating records, or moving files between applications. AI-driven automation works differently. It can classify information, interpret language, detect anomalies, predict outcomes, and decide the most appropriate next steps based on context.

The difference can best be summarized like this:

Traditional automation and RPA AI-driven automation
Follows explicit rules Uses models and data patterns
Works well with structured inputs Can work with unstructured inputs such as text, images, or documents
Repeats known steps Can adapt to variation
Is sensitive to interface or process changes Can recommend, classify, or prioritize
Usually automates task execution Supports decision-making as well as execution
Depends on humans to define every rule Requires monitoring, evaluation, and governance

The strongest approach is often a combination. RPA can execute predictable actions, whereas AI can interpret the situation, enrich the data, make a recommendation, or decide whether human review is needed.

Why AI automation is gaining attention 

AI automation is becoming more relevant because organizations are under pressure to improve productivity while managing growing data complexity. Many business processes now involve large volumes of text, images, transactions, sensor data, customer interactions, and operational signals. Traditional automation struggles when inputs are inconsistent. AI can help with that complexity.

Current market research also shows both opportunity and caution. AI adoption is widespread, but many organizations still struggle to embed AI into core workflows. Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value, or inadequate risk controls (Source: Gartner, “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027,” press release, June 25, 2025). That is an important signal for decision-makers: AI automation should not start with technology hype; it should start with a clear process, a measurable outcome, available data, and clear governance.

How AI automation works

AI automation typically combines several components

  1. Data ingestion 

    The workflow starts with data from systems, files, documents, APIs, emails, applications, sensors, or customer interactions. This data may be structured, semi-structured, or unstructured.

    A platform like Databricks can ingest and process these sources in batch or streaming workflows. This matters because AI automation needs access to reliable and current data.

  2. Data preparation

    Raw data is rarely ready for automation; it needs cleaning, validation, deduplication, transformation, and enrichment.

    For example, customer messages may need language detection, document data may need extraction, and transaction records may need to be matched with master data.

  3. AI model or agent logic

    The AI layer interprets the data and supports the next step. This may involve:

    • classifying a customer request

    • extracting information from a document

    • predicting risk

    • detecting unusual behavior

    • summarizing a case

    • recommending an action

    • selecting a tool or workflow step

    In more advanced workflows, AI agents can call tools, search knowledge sources, query databases, or trigger actions under defined controls.

  4. Workflow orchestration

    Automation needs coordination. The system must know which steps run first, which checks are required, when to involve a human, and what should happen after a decision.

    Databricks Workflows can support this type of orchestration by running data and AI tasks in a controlled sequence.

  5.  Human review

    AI automation does not mean every decision should be fully autonomous. Many business processes need human approval, especially when financial, legal, operational, or customer impact is high.

    A good AI workflow defines clear thresholds. Low-risk tasks can be handled automatically. Medium-risk cases can be recommended to a reviewer. High-risk cases should be escalated.

  6. Monitoring and improvement

    AI automation must be monitored. Teams need to track model performance, errors, data drift, cost, latency, usage, and business outcomes.

    A lack of monitoring makes AI unreliable, whereas proper monitoring enables teams to improve systems over time.

Where Databricks fits in AI automation

Each of the six components above touches Databricks in some way: ingestion, preparation, model logic, orchestration, and monitoring all run as Databricks capabilities rather than six separate tools bolted together. That consolidation is the reason Databricks is well-suited to AI automation. A chatbot or workflow tool alone is not enough when an organization needs to automate complex business processes that cross multiple systems. This is especially important when AI automation crosses those systems. For example, an automated claims workflow may need policy data, customer history, documents, fraud signals, geospatial data, and decision rules. Databricks can bring these data sources together and make them usable for AI models and workflow logic. Bringing the data together is only half the equation, though. What gets built on top of it, the actual AI logic, monitoring, and human oversight- is a discipline of its own called AI Engineering.

Examples of AI engineering in automation

AI engineering includes designing, building, and maintaining AI systems in production. It goes beyond a model or prompt. It includes data pipelines, evaluation, monitoring, security, integration, and human oversight.

AI Engineering can be used for:

  • Intelligent document processing. AI can extract fields from contracts, invoices, forms, or reports. A workflow can validate those fields against internal systems, flag inconsistencies, and route exceptions to humans. Databricks can support the underlying data preparation, quality checks, model evaluation, and storage of structured outputs.
  • Customer service triage. AI can classify incoming messages, detect urgency, summarize context, and suggest responses. The system can route simple requests automatically and escalate complex cases. This type of workflow requires integration with customer data, knowledge bases, conversation history, and performance monitoring.
  • Predictive maintenance. AI can analyze sensor data, maintenance logs, and operational signals to predict failure risks. Automation can then create alerts, prioritize inspections, or recommend spare parts. This requires reliable streaming or batch pipelines, feature engineering, model monitoring, and integration with operational systems.
  • Finance and compliance review. AI can detect anomalies in transactions, identify policy deviations, or summarize documents for review. Automation can prioritize cases based on risk and evidence. Human review remains important, but AI helps reduce manual screening and improves focus.
  • Marketing automation with AI. AI can segment audiences, predict next-best actions, personalize content, and score leads. Automation can then trigger campaigns or recommendations. The challenge is governance: teams need to control data usage, consent, model quality, and measurement.

Benefits of AI automation

AI automation can create value for organizations in several ways.

Some examples of this are:

  • Faster process execution. AI can reduce manual steps in workflows that involve reading, classifying, checking, or summarizing information. This can shorten cycle times in operations, customer service, finance, and IT.
  • Better handling of variation. Traditional automation often fails when inputs differ from the expected format. AI can handle more variation, especially in text, documents, and behavioral data.
  • Improved decision support. AI can help prioritize cases, detect risk, recommend actions, or surface relevant context. This supports better decisions without requiring teams to manually inspect every record.
  • Scalable knowledge work. Many office processes involve repetitive knowledge work: reading documents, comparing information, writing summaries, and checking exceptions. AI automation helps scale that work while keeping human review for judgment-heavy decisions.
  • Stronger process insight. When AI workflows are built on a data platform, organizations can measure where processes slow down, where exceptions occur, and which decisions lead to better outcomes. 

Implementation steps for AI automation

Step 1: Select the right process. Start with a process that has enough volume, clear pain points, and measurable value. Good candidates include document processing, support triage, reporting preparation, risk scoring, data quality checks, and operational alerts. Avoid starting with a process that is poorly understood or politically sensitive.

Step 2: Define success criteria. Success should be measurable. Examples include reduced handling time, fewer manual checks, better data quality, faster response times, higher first-time-right rates, or improved prioritization.

Step 3: Assess data readiness. AI automation depends on data. Check whether the required data is available, accurate, governed, and accessible. If the data is fragmented or unreliable, solve that first.

Step 4: Design the workflow. Map the process steps, AI decisions, business rules, human review points, and system integrations. Define where AI recommends, where it acts, and where humans approve.

Step 5: Build a proof of concept. A PoC tests the real workflow with representative data. It should not be limited to a demo prompt. The goal is to test feasibility, value, risks, and operational requirements.

Step 6: Move to production with controls. Production AI automation needs monitoring, access control, cost control, model evaluation, logging, documentation, and rollback options.

Common mistakes in AI automation

Even well-planned AI automation projects tend to run into the same handful of problems. These problems are worth double-checking before committing to a budget and plan. The five most common mistakes are:

  • Automating a broken process. AI should not be used to hide unclear ownership, poor data quality, or inconsistent business rules. Start by fixing the process design first.
  • Starting too broad. Large transformation programs often move slowly. Start with a focused use case that can prove value and teach the organization what is needed.
  • Treating AI as fully autonomous too early. AI can support decisions, but many workflows need human oversight. Autonomy should increase only when performance, risk, and governance are understood.
  • Ignoring governance. AI workflows often use sensitive data. Access, lineage, auditability, and model behavior should be controlled from the start.
  • Measuring only technical success. A working model is not enough. Measure business impact: time saved, quality improved, risk reduced, or revenue protected.

RevoData’s approach to AI automation

RevoData helps organizations move from rule-based automation to AI-driven workflows on Databricks. The approach is practical; together we identify the right process, assess data readiness, design a controlled architecture, and build a PoC that can scale.

As a Databricks Gold Partner with 100% Databricks-certified consultants, RevoData brings the technical depth needed for production AI systems. The team combines data engineering, AI engineering, governance, and platform knowledge to get results. RevoData also has a strong culture of diligence and constant development, with one of the highest numbers of Databricks Champions in EMEA.

For clients, this means assured quality and capability. Your AI automation is not treated as a standalone tool project. It becomes part of a governed data and AI platform.

Ready to explore AI automation with Databricks? RevoData can help you select the right use case, build a proof of concept, and design the path to production. 

Final thought

AI automation does more than speed up traditional automation. It changes what can be automated in the first place. Rule-based automation executes known steps. AI-driven automation can interpret information, learn from data, recommend actions, and support more complex workflows. The organizations that benefit most are the ones that combine AI with reliable data, clear governance, and practical process design. Databricks provides the platform foundation for that shift. RevoData helps organizations turn that foundation into working AI automation: focused use cases, certified expertise, scalable architecture, and a clear path from PoC to production.

FAQ's

AI automation is the use of artificial intelligence to automate tasks, decisions, or workflows that require interpretation, prediction, or contextual understanding. It can include document processing, classification, recommendations, anomaly detection, summarization, and AI agents.

Start with a business process that has measurable value, enough volume, and available data. Define success criteria, assess data quality, design the workflow, and test the approach with a focused proof of concept.

AI automation combines data pipelines, AI models or agents, workflow orchestration, business rules, human review, and monitoring. The AI layer interprets information or recommends actions, while the automation layer executes the workflow under defined controls.

RPA automates repetitive, rule-based tasks by following predefined steps. AI automation uses models and data to interpret inputs, detect patterns, make predictions, and support decisions. RPA is best for stable processes; AI automation is better for variable, data-rich, or context-heavy workflows.

Benefits include faster processing, better handling of unstructured data, improved prioritization, fewer manual checks, stronger decision support, and better process visibility. The value is highest when AI automation is connected to reliable data and governance.

Databricks provides a platform for data engineering, machine learning, generative AI, governance, and workflow orchestration. This makes it suitable for complex AI automation where data from multiple sources must be processed, governed, and used in production workflows.

]]>
RevoData in Het Financieele Dagblad https://revodata.nl/revodata-in-het-financieele-dagblad/ Wed, 29 Apr 2026 08:15:32 +0000 https://revodata.nl/?p=6948

We were recently featured in Het Financieele Dagblad (FD), where our CEO,Ralph Kootker, addressed the critical challenges of scaling AI. In a landscape dominated by global tech giants, Ralph highlights how a best-of-both-worlds approach, leveraging the power of the public cloud while maintaining data sovereignty through open-source standards, is the key to strategic independence.

While the original discussion took place in Dutch, the challenges of data portability and governance are universal. Below is the translated version of the full interview.

AI Requires a Unified Data Platform

How do you quickly extract value from AI technologies when your data is scattered across different systems, tools, and cloud platforms? Without a coherent data architecture, initiatives often get stuck in the experimental phase. An integrated platform where data, analytics, and AI converge within a single ecosystem is a crucial prerequisite for this.

This allows applications to be developed consistently and achieve impact on a larger scale. Ralph Kootker, founder and CEO of RevoData, says that exactly this balance is central to current AI development. The Amsterdam-based data and AI consultancy helps organizations set up their data architecture so that AI applications remain scalable and manageable. “The core question is not only what you can do with AI, but especially how you maintain control over your data and processes. That is essential to realize value for your company.”

Greater Depth

RevoData was founded in 2022 and is part of the Atomic Group, which also includes cloud providers Uniserver and CloudNation. The company specializes fully in the data platform Databricks. According to Kootker, this focus was a conscious choice. “We work with one technology vendor because it allows you to build true depth and expertise,” he says. “Databricks is based on open-source components and standards. That makes it easier to remain strategically independent, even if you run on the public cloud.”

This open architecture is becoming increasingly important as organizations look more critically at their dependence on American tech companies. Kootker states that the importance of data sovereignty is especially relevant for public institutions and sectors where privacy is central. “Our economy runs largely on American technology. Many organizations realize they have their backs against the wall if something changes geopolitically.”

Scale and Computing Power

At the same time, the public cloud remains essential for innovation for the time being. “The scale and computing power of it are still unparalleled,” he says. “But you must ensure that your architecture remains portable, so that you can move workloads to a private cloud if necessary.” RevoData works on this with Dutch infrastructure partners, such as sister company Uniserver. By running Databricks’ open-source components in private cloud environments, a hybrid model is created that, according to Kootker, offers the best of both worlds: strategically independent, but tactically flexible.

Speed

In addition to data sovereignty, he sees another challenge: the speed at which AI tools can generate software. “We are soon going to have engineers who produce enormous amounts of code with the help of AI, without checking it properly. If companies do not set up clear processes and quality safeguards, things will go wrong eventually.”

Therefore, the emphasis at RevoData is not just on technology, but primarily on processes and architecture. Consultants help organizations actually bring AI models to production and manage them safely. “Everyone can build an AI model nowadays,” says Kootker. “But making it truly operational and maintaining it is a completely different discipline. That is only possible if your architecture, data, and governance are set up correctly from the start.”

Diverse Team

The composition of their team also plays a role in this. The company consciously works with experienced specialists from different countries and backgrounds. “Diversity and seniority make the difference. Different perspectives almost always lead to better solutions.”

Is your data architecture ready for the next step?

Most AI projects fail not because of the math, but because of the foundation. Ensure your organization is built for scale, portability, and sovereignty.

Book a complimentary 30-minute Data & AI Audit with our team to evaluate your current architecture.

]]>
RevoData achieves Gold Databricks Partner status https://revodata.nl/revodata-achieves-gold-databricks-partner-status/ https://revodata.nl/revodata-achieves-gold-databricks-partner-status/#respond Mon, 09 Feb 2026 10:17:22 +0000 https://revodata.nl/?p=6594

Amsterdam, February 2026 – RevoData, the Amsterdam-based data and AI consultancy, has achieved Databricks Gold Partner status (formerly named Elite). This achievement highlights the extensive collaboration between the two companies and acknowledges RevoData’s deep expertise in building high-quality solutions on the Databricks platform.

“Becoming a Gold Databricks Partner has been a long-term goal for RevoData. To achieve this is truly an honor and only possible through the dedication and effort of our team. We look forward to further deepening our collaboration with Databricks and building revolutionary solutions for our customers.” – Ralph Kootker, CEO of RevoData


RevoData rose from Registered Partner to Gold Partner in four years through innovative solutions, a dedication to growth, and continued value delivery. This combination resulted in one of its geospatial solutions, used for flood prevention by applying AI on drone imagery, to win a recognized industry award for innovation. RevoData’s further achievements include: 

  • Six Databricks Champions, the highest number in the Netherlands 
  • Providing training for a wide range of companies through the Databricks Training Program 
  • Having all consultants 100% certified in their field of expertise

 

Looking ahead, RevoData will continue to strengthen its core services[RK8] across the Nordics and Benelux while doubling down on its expertise in geospatial solutions and LLMOps.

“Geospatial and LLMOps are two areas in which I believe our knowledge can serve the market. Through our deep expertise and focus on Databricks, we take companies beyond the hype, elongating model lifecycles while increasing value delivery. Where AI is already booming, the use cases for geospatial data are slowly but steadily growing. With our in-house experts, we bridge the knowledge gap to bring use cases in production” – Ralph Kootker, CEO at RevoData 

 

 

About RevoData 

RevoData is a Dutch consultancy firm providing solutions that accelerate innovation, are user-friendly, and are built on Databricks. Through its breadth of expertise, the company builds sustainable solutions that drive business value. RevoData provides training and guidance to ensure companies can effectively use and maintain delivered projects. 

]]>
https://revodata.nl/revodata-achieves-gold-databricks-partner-status/feed/ 0
Eyes to the Sky: LiDAR Point Cloud in Databricks for Urban Canopy Insights https://revodata.nl/eyes-to-the-sky-lidar-point-cloud-in-databricks-for-urban-canopy-insights/ https://revodata.nl/eyes-to-the-sky-lidar-point-cloud-in-databricks-for-urban-canopy-insights/#respond Fri, 23 Jan 2026 14:32:43 +0000 https://revodata.nl/?p=6563

Editor’s note: This post was originally published June 20th, 2025. 

What is the Sky View Factor and why does it matter in our cities?

I’m back with another sunny, summer-inspired geospatial adventure!
Ever wonder how much sky you can actually see when you’re standing on a busy city street or under a cluster of city trees? That visible slice of sky called the Sky View Factor (SVF) has a surprisingly big say in how cities heat up, cool down, and even how comfortable we feel outside. The less sky you see, the more heat gets trapped between buildings and trees, creating those urban heat islands, where city temperatures climb higher than surrounding areas. These heat islands don’t just make summer days unbearable, they can worsen air pollution, increase energy demand for cooling, strain public health by amplifying heat-related illnesses, and even accelerate the wear and tear on city infrastructure.

Measuring SVF takes more than just a weather app or a quick satellite snapshot, you need a detailed, 3D, street-level view of the city. That’s where LiDAR point cloud data shines, capturing billions of laser-scanned points from every rooftop, treetop, street, and sidewalk. This treasure trove of 3D data lets us model the urban canopy, the intricate layer of buildings and greenery that controls how sunlight and air flow through the cityscape. From this, we can calculate SVF and generate cool fisheye plots that show exactly how much sky you’d see lying anywhere on the ground.

Back in my master’s program, some classmates and I built an app for the municipality of The Hague that let users click on a map or upload a list of points to get SVF values and how much the sky is blocked by buildings and trees. You can check out our full report with the front-end and backend code here. Back then, we ran the calculations using NumPy arrays and plenty of good old-fashioned for-loops. I’m now reworking the project using PySpark, with a focus on scalable, data warehousing for big data analytics rather than real-time, on-the-fly processing, without delving deeply into the underlying mathematical computations. That’s where Databricks shines: its cloud-native platform effortlessly handles massive LiDAR datasets, turning mountains of 3D points into quick, actionable insights. Whether you’re a city planner aiming to cool down urban streets or simply a curious urban data explorer, it’s never been easier or more fun to look up and ask: how much sky do we really see?

Implementation:

The following implementation is part of the training we offer at RevoData focused specifically on leveraging the geospatial capabilities of Databricks.

In this implementation, I want to focus on code refactoring and explore the options available to start with the low-hanging fruit, highlighting which parts of the code can be adapted for distributed processing with minimal changes.

Let’s get started!

Datasets

As I mentioned earlier, we originally used point cloud data from the City of The Hague for this project. But today, I’m taking you to Washington, partly to switch up the scenery for myself, and partly so you can download the data more easily from an English-language website. The data is fully available here. I also generated a grid of 576 points across the area, which we’ll use to calculate the Sky View Factor (SVF). 

Import libraries

In this implementation, we need couple of libraries which I import in one go:

				
					import math
import numpy as np
import boto3
import os
import matplotlib.pyplot as plt
import pdal
import json
import io
import pyarrow as pa
from pyspark.sql.functions import col, sqrt, pow, lit, when, atan2, degrees, floor
from pyspark.sql.types import StructType, StructField, DoubleType, FloatType, IntegerType, ShortType, LongType, ByteType, BooleanType, MapType, StringType, ArrayType
import pandas as pd
from sedona.spark import *
from pyspark.sql import functions as F
from pyspark.sql.window import Window
import base64
from PIL import Image
				
			

Data ingestion

For the generated grid points, Apache Sedona makes geospatial data ingestion remarkably easy.

				
					
config = SedonaContext.builder() .\
    config('spark.jars.packages',
           'org.apache.sedona:sedona-spark-shaded-3.3_2.12:1.7.1,'
           'org.datasyslab:geotools-wrapper:1.7.1-28.5'). \
    getOrCreate()

sedona = SedonaContext.create(config)

# The path to the grid geopackage and point cloud las file
dataset_bucket_name = "revodata-databricks-geospatial"
dataset_input_dir="geospatial-dataset/point-cloud/washington"
gpkg_file = "grid/pc_grid.gpkg"
pointcloud_file = "las-laz/1816.las"
input_path = f"s3://{dataset_bucket_name}/{dataset_input_dir}/{pointcloud_file}"

# Read the grid data
df_grid = sedona.read.format("geopackage").option("tableName", "grid").load(f"s3://{dataset_bucket_name}/{dataset_input_dir}/{gpkg_file}").withColumnRenamed("geom", "geometry").withColumn("x1", F.expr("ST_X(geometry)")).withColumn("y1", F.expr("ST_Y(geometry)")).select("fid", "x1", "y1", "geometry")

num_partitions = math.ceil(df_grid.count()/2)

				
			

To ingest point cloud data, we can use libraries like laspy or PDAL. In this case, I used PDAL, applying a few read-time optimizations to efficiently convert the output array into a PySpark DataFrame:

				
					
def _create_arrow_schema_from_pdal(pdal_array):
    """Create Arrow schema from PDAL array structure."""
    fields = []
    
    # Map PDAL types to Arrow types
    type_mapping = {
        'float32': pa.float32(),
        'float64': pa.float64(),
        'int32': pa.int32(),
        'int16': pa.int16(),
        'uint8': pa.uint8(),
        'uint16': pa.uint16(),
        'uint32': pa.uint32()
    }
    
    for field_name in pdal_array.dtype.names:
        field_type = pdal_array[field_name].dtype
        arrow_type = type_mapping.get(str(field_type), pa.float32())  # default to float32
        fields.append((field_name, arrow_type))
    
    return pa.schema(fields)

def _create_spark_schema(arrow_schema):
    """Convert PyArrow schema to Spark DataFrame schema."""
    spark_fields = []
    
    type_mapping = {
        pa.float32(): FloatType(),
        pa.float64(): DoubleType(),
        pa.int32(): IntegerType(),
        pa.int16(): ShortType(),
        pa.int8(): ByteType(),
        pa.uint8(): ByteType(),
        pa.uint16(): IntegerType(),  # Spark doesn't have unsigned types
        pa.uint32(): LongType(),     # Spark doesn't have unsigned types
        pa.string(): StringType(),
        # Add other type mappings as needed
    }
    
    for field in arrow_schema:
        arrow_type = field.type
        spark_type = type_mapping.get(arrow_type, StringType())  # default to StringType
        spark_fields.append(
            StructField(field.name, spark_type, nullable=True)
        )
    
    return StructType(spark_fields)


def pdal_to_spark_dataframe_large(pipeline_config, spark, chunk_size=1000000):
    """Streaming version for very large files."""
    pipeline = pdal.Pipeline(json.dumps(pipeline_config))
    pipeline.execute()
    
    # Get schema from first array
    first_array = pipeline.arrays[0]
    schema = _create_arrow_schema_from_pdal(first_array)
    
    # Create empty RDD
    rdd = spark.sparkContext.emptyRDD()

    
    # Process arrays in chunks
    for array in pipeline.arrays:
        for i in range(0, len(array), chunk_size):
            chunk = array[i:i+chunk_size]
            data_dict = {name: chunk[name] for name in chunk.dtype.names}
            arrow_table = pa.Table.from_pydict(data_dict, schema=schema)
            pdf = arrow_table.to_pandas()
            chunk_rdd = spark.sparkContext.parallelize(pdf.to_dict('records'))
            rdd = rdd.union(chunk_rdd)
    
    # Convert to DataFrame
    return spark.createDataFrame(rdd, schema=_create_spark_schema(schema))
				
			
				
					
pipeline_config = {
    "pipeline": [
        {
            "type": "readers.las",
            "filename": input_path,
        }
    ]
}

# Convert point cloud array to Spark DataFrame
df_pc = pdal_to_spark_dataframe_large(pipeline_config, spark)
df_pc = df_pc.withColumn("geometry", F.expr("ST_Point(X, Y)"))
df_pc.write.mode("overwrite").saveAsTable(f"geospatial.pointcloud.wasahington_pc")
				
			

Identifying point cloud data surrounding each grid point

Next, we need to retrieve all point cloud data within 100 meters for buildings and high vegetation, these are used for SVF calculation. For ground points, we only consider those within 10 meters, as they’re used solely to estimate the elevation of each grid point.

				
					df_selected = df_pc.select("X", "Y", "Z", "Classification")

dome_radius = 100
height_radius = 10

# Register as temp views
df_pc.createOrReplaceTempView("pc_vw")
df_grid.createOrReplaceTempView("grid_vw")

# Perform spatial join using ST_DWithin with 100 meters
grid_join_pc = spark.sql(f"""
    SELECT 
        g.fid, 
        ST_X(g.geometry) AS x1,
        ST_Y(g.geometry) AS y1,
        p.classification,
        p.x AS pc_x,
        p.y AS pc_y,
        p.z AS pc_z,
        ST_Distance(g.geometry, p.geometry) AS distance,
        g.geometry AS g_geometry,
        p.geometry AS pc_geometry 
    FROM grid_vw g
    JOIN pc_vw p
        ON ST_DWithin(g.geometry, p.geometry, {dome_radius})
    WHERE p.classification IN (5, 6) OR (p.classification = 2 AND ST_DWithin(g.geometry, p.geometry, {height_radius}))
""")
				
			

Estimating grid point elevation from nearby point cloud data (10m Radius)

Here, we determine the elevation of each grid point by identifying its dominant surrounding class, either building or ground, and then computing the average elevation of that class within a defined radius.

 

				
					# Filter only classification 2 and 6 and count occurrences of (fid, classification)
grouped = grid_join_pc.filter(
    (F.col("classification").isin(2, 6)) & (F.col("distance") <= height_radius)
).groupBy("fid", "classification").count()

# Define window: partition by fid, order by count descending
window_spec = Window.partitionBy("fid").orderBy(F.desc("count"))

# Apply row_number
ranked = grouped.withColumn("rn", F.row_number().over(window_spec))

# Compute the average elevation for each grid point using nearby point cloud data within a specified radius.
grid_pc_elevation = grid_join_pc.join(g_classification_df, on=["fid", "classification"]).filter(
    (F.col("distance") <= height_radius)
).groupBy("fid").agg(
    (F.sum("pc_z") / F.count("pc_z")).alias("height")
)

# Combine point cloud data with classification info and computed height, optimized with repartitioning.
grid_pc_elevation_all = grid_join_pc.withColumnRenamed("classification", "p_classification").join(g_classification_df, on=["fid"]).join(grid_pc_elevation, on=["fid"]).repartitionByRange(num_partitions, "fid")

# Filter out ground points (e.g., class 2) to retain only buildings and high vegetation points for analysis.
grid_pc_cleaned = grid_pc_elevation_all.filter("p_classification != 2").repartitionByRange(num_partitions, "fid")
				
			

Creating the dome, generating the plot, and calculating the SVF

The dome is a representation of the sky, going from the horizon all the way to the zenith (directly on top) of the viewpoint. The dome can be split into sectors based on horizontal and vertical directions, in essence creating a dome-like shaped grid. The units used to split the sectors are 2 degrees horizontally (azimuth angle), and 1 degree vertically (elevation angle), which are considered as appropriate values for calculation.

To calculate the Sky View Factor (SVF), point cloud data is projected onto a dome divided into sectors, marking which sectors are blocked from view. The closest point in each sector determines the obstruction, and if that point is a building, all sectors below it in that direction are also considered blocked. The unobstructed proportion of the dome’s area gives the SVF. For clarity, the results are visualized in a circular plot showing which sectors are clear sky or obstructed by buildings or vegetation, oriented to the north for easy interpretation.

				
					# Calculate raw azimuth angle (in degrees) from each grid point to each point in the point cloud.
# Shifted by -90 to align with the 0° direction being north.
grid_pc_az = grid_pc_cleaned.withColumn(
    "azimuth_raw",
    degrees(F.atan2(F.col("pc_y") - F.col("y1"), F.col("pc_x") - F.col("x1"))) - 90
)

# Normalize azimuth angle to fall within the range [0, 360).
grid_pc_az = grid_pc_az.withColumn(
    "azimuth",
    when(F.col("azimuth_raw") < 0, F.col("azimuth_raw") + 360).otherwise(F.col("azimuth_raw"))
)

# Drop the temporary azimuth_raw column to clean up the DataFrame.
grid_pc_az = grid_pc_az.drop("azimuth_raw")

# Calculate the elevation angle (in degrees) from the grid point to each point in the point cloud.
# Height is divided by 1000 to convert from millimeters to meters if necessary.
grid_pc_az = grid_pc_az.withColumn(
    "elevation",
    degrees(F.atan2(F.col("pc_z") - F.col("height") / 1000, F.col("distance")))
)

# Bin azimuth angles into 2-degree intervals (0–179 bins for 360°).
grid_pc_az = grid_pc_az.withColumn("azimuth_bin", F.floor(F.col("azimuth") / 2))

# Get the minimum elevation angle across all records to define the lower bound of elevation bins.
min_val = F.lit(grid_pc_az.select(F.min("elevation")).first()[0])

# Get the maximum elevation angle across all records to define the upper bound of elevation bins.
max_val = F.lit(grid_pc_az.select(F.max("elevation")).first()[0])

# Compute bin width by dividing elevation range into 89 equal parts (90 bins total).
bin_width = (max_val - min_val) / 89

# Bin elevation angles into 90 intervals, ensuring they stay within the [0, 89] range.
grid_pc_az = grid_pc_az.withColumn("elevation_bin", 
    F.least(
        F.greatest(
            F.floor(
                (F.col("elevation") - F.lit(min_val)) / 
                F.lit((max_val - min_val)/90)
            ).cast("int"),
            F.lit(0)  # Clamp minimum bin index to 0
        ),
        F.lit(89)  # Clamp maximum bin index to 89
    )
)


# Define a window that partitions the data by azimuth and elevation bins,
# and orders points within each bin by their distance to the grid point.
window_spec = Window.partitionBy("azimuth_bin", "elevation_bin").orderBy("distance")

# Assign a row number within each azimuth-elevation bin, so the closest point (smallest distance) gets rank 1.
df_with_rank = grid_pc_az.withColumn("rn", F.row_number().over(window_spec))

# Keep only the closest point (rank 1) in each bin and drop the temporary rank column.
# Then repartition the result by 'fid' to optimize parallel processing in subsequent steps.
closest_points = df_with_rank.filter(col("rn") == 1).drop("rn").repartitionByRange(num_partitions, "fid")
				
			
				
					
def create_dome(pdf: pd.DataFrame, max_azimuth: int = 180, max_elevation: int = 90) -> np.ndarray:
    """
    Creates a dome matrix based on azimuth and elevation bins, with obstruction handling for buildings.
    """
    dome = np.zeros((max_azimuth, max_elevation), dtype=int)
    domeDists = np.zeros((max_azimuth, max_elevation), dtype=float)

    for _, row in pdf.iterrows():
        a = int(row["azimuth_bin"])
        e = int(row["elevation_bin"])
        dome[a, e] = row["p_classification"]
        domeDists[a, e] = row["distance"]

    # Mark parts of the dome that are obstructed by buildings
    if np.any(dome == 6):  # 6 = buildings
        bhor, bver = np.where(dome == 6)
        builds = np.stack((bhor, bver), axis=-1)
        shape = (builds.shape[0] + 1, builds.shape[1])
        builds = np.append(builds, (bhor[0], bver[0])).reshape(shape)
        azimuth_change = builds[:, 0][:-1] != builds[:, 0][1:]
        keep = np.where(azimuth_change)
        roof_rows, roof_cols = builds[keep][:, 0], builds[keep][:, 1]
        for roof_row, roof_col in zip(roof_rows, roof_cols):
            condition = np.where(np.logical_or(
                domeDists[roof_row, :roof_col] > domeDists[roof_row, roof_col],
                dome[roof_row, :roof_col] == 0
            ))
            dome[roof_row, :roof_col][condition] = 6

    return dome

# Plot dome
def generate_plot_image(dome):
    # Create circular grid
    theta = np.linspace(0, 2*np.pi, 180, endpoint=False)
    radius = np.linspace(0, 90, 90)
    theta_grid, radius_grid = np.meshgrid(theta, radius)

    Z = dome.copy().astype(float)
    
    Z = Z.T[::-1, :]  # Transpose and flip vertically

    Z[Z == 0] = 0
    Z[np.isin(Z, [5])] = 0.5
    Z[Z == 6] = 1

    if Z[Z == 6].size == 0:
        Z[0, 0] = 1  # Force plot to show something

    fig = plt.figure(figsize=(4, 4))
    ax = fig.add_subplot(111, projection='polar')
    cmap = plt.get_cmap('tab20c')
    ax.pcolormesh(theta, radius, Z, cmap=cmap)
    ax.set_ylim([0, 90])
    ax.tick_params(labelleft=False)
    ax.set_theta_zero_location("N")
    ax.set_xticks([])
    ax.set_yticks([])

    buf = io.BytesIO()
    plt.savefig(buf, format='png', bbox_inches='tight', pad_inches=0)
    plt.close(fig)
    buf.seek(0)
    img_base64 = base64.b64encode(buf.read()).decode('utf-8')
    return img_base64


def process_and_plot(pdf: pd.DataFrame) -> pd.DataFrame:
    fid = pdf["fid"].iloc[0]

    # Create dome with building/vegetation obstruction
    dome = create_dome(pdf)

    # Generate base64-encoded fisheye plot image
    plot_base64 = generate_plot_image(dome)

    # Compute SVF and obstruction metrics
    SVF, tree_percentage, build_percentage = calculate_SVF(100, dome)

    return pd.DataFrame(
        [[fid, dome.tolist(), plot_base64, SVF, tree_percentage, build_percentage]],
        columns=["fid", "dome", "plot", "SVF", "treeObstruction", "buildObstruction"]
    )
				
			

Results

As you can see in the code above, some parts still rely on NumPy structures for operations like dome construction, plotting and SVF calculation. Since each grid point corresponds to a single dome, we can safely apply these NumPy-based functions on a per-row basis. To do this efficiently in PySpark , without too much code refactoring, I use the applyInPandas method. This allows us to apply our existing Pandas-based logic directly to each group of rows (grouped by the "fid" column) within the PySpark DataFrame. This way, we can leverage distributed processing in Spark while reusing existing, well-tested NumPy code for the dome and SVF calculations.

				
					# Desired schema
output_schema = StructType([
    StructField("fid", IntegerType()),
    StructField("dome", ArrayType(ArrayType(IntegerType()))),
    StructField("plot", StringType()),
    StructField("SVF", FloatType()),
    StructField("treeObstruction", FloatType()),
    StructField("buildObstruction", FloatType())
])

result_df = closest_points.groupBy("fid").applyInPandas(process_and_plot, schema=output_schema)
result_df.write.mode("overwrite").saveAsTable(f"geospatial.pointcloud.wasahington_grid")
				
			
				
					
# Fetch the the grid point with fid = 105 for a sample visualization
pdf = result_df.filter("fid = 105").select("fid", "plot").toPandas()

for index, row in pdf.iterrows():
  # Decode base64 string to bytes
  img_bytes = base64.b64decode(img_base64)

  # Load image with PIL
  image = Image.open(io.BytesIO(img_bytes))

  # Display using matplotlib (preserves original colors)
  plt.figure(figsize=(6, 6))
  plt.imshow(image)
  plt.axis('off')  # Hide axes
  plt.show()
				
			

What is next?

In our live training Databricks Geospatial in a Day at RevoData Office, we’ll delve deeper into the logic behind this code and use this example to demonstrate how to:

  • Set up a cluster capable of processing point cloud data
  • Visualize LiDAR point clouds directly in Databricks
  • Efficiently partition point cloud data for distributed processing
  • Tackle the challenges of code migration and minimal refactoring

Go ahead and grab your spot for the training using the link below, can’t wait to see you there!
https://revodata.nl/databricks-geospatial-in-a-day/

Picture of Melika Sajadian

Melika Sajadian

Senior Geospatial Consultant at RevoData, sharing with you her knowledge about Databricks Geospatial

]]>
https://revodata.nl/eyes-to-the-sky-lidar-point-cloud-in-databricks-for-urban-canopy-insights/feed/ 0
Building a Geospatial Time Machine https://revodata.nl/building-a-geospatial-time-machine/ https://revodata.nl/building-a-geospatial-time-machine/#respond Fri, 23 Jan 2026 13:46:16 +0000 https://revodata.nl/?p=6551

Editor’s note: This post was originally published June 12th, 2025. 

What is environmental change detection?

Think of it as a time machine for planet Earth. No flux capacitor, no DeLorean, and no Marty McFly needed!😉 Just some cool tech that lets us peek into how our world changes over time.

Detecting and understanding environmental changes is essential for informed decision-making, sustainable development, and effective disaster risk reduction. By systematically identifying and monitoring trends such as deforestation, urban expansion, and coastal erosion at an early stage, policymakers, urban planners, environmental agencies, and other relevant stakeholders can implement timely and proactive measures to mitigate adverse impacts on ecosystems, biodiversity, and vulnerable human populations.

Change detection provides a critical evidence base for assessing the effectiveness of existing environmental policies and conservation strategies. It also plays a key role in informing the planning and development of resilient infrastructure capable of withstanding future environmental stresses. Furthermore, by offering accurate and up-to-date information, it supports more efficient resource allocation and helps prioritize areas facing the greatest risks.

Above all, the ability to detect and interpret environmental changes enhances society’s capacity to respond to climate-related challenges such as sea-level rise, extreme weather events, and habitat degradation. In this context, change detection serves as a vital tool for promoting more responsible and adaptive management of the planet in an era marked by rapid and often unpredictable transformation. While this “time machine” can’t rewrite history, it empowers us to learn from the past and chart a course toward a more sustainable future.

For today’s adventure, I’m using Databricks and Apache Sedona , they’re basically the power tools for working with geospatial data without your computer having a meltdown. And where are we headed? San Francisco,specifically the SoMa (South of Market) area! We’re gonna see how it looked in August 2022 versus February 2025. Trust me, even in just a couple years, you’d be surprised how much a neighborhood can transform. It’s like watching your city grow up in fast-forward.

Implementation

The following implementation is part of the training we offer at RevoData focused specifically on leveraging the geospatial capabilities of Databricks.

Remote sensing and photogrammetry cover a lot of ground, you’ve got active and passive sensors, different types of resolution (spatial, spectral, temporal, radiometric), geo-referencing, orthophoto generation, and electromagnetic spectrum analysis. Instead of getting bogged down in all the technical details, I’m going to focus on why I chose certain methods and what other options could potentially be considered. Let’s get started.

Datasets

First things first, we need data, but what kind? For change detection in an urban environment, we typically rely on satellite or aerial images captured in at least four spectral bands: Red, Green, Blue, and Near Infrared (NIR). Why NIR? Well, it all comes down to physics. Different sensors detect different ‘flavors’ of energy, from visible light our eyes see to invisible infrared or microwave radiation. Each sensor type reveals unique information about Earth’s surface based on the specific energy waves it can measure. What the sensor ‘sees’ depends on how objects interact with these waves:

  • Absorption: The surface soaks up the energy (like dark pavement heating in sunlight)
  • Reflection: The energy bounces back (like light mirroring off a lake)
  • Transmission: The energy passes through (like sunlight through clear water)

Different materials, such as water, soil, concrete, vegetation, each have unique ‘fingerprints’ in how they handle these waves. That’s how we can identify and monitor features from satellites or aircraft!

These interactions allow us to calculate indices such as the Normalized Difference Vegetation Index (NDVI) and the Normalized Difference Water Index (NDWI). These are simple and useful for classifying each image pixel as vegetation, water, or bare ground. You can also go a step further and apply machine learning or pattern recognition algorithms, like the maximum likelihood classifier, to identify different phenomena across the landscape.

The good news? This kind of imagery is freely available. You can easily access the data in GeoTiff format through the NOAA Data Access Viewer and begin your own journey through space and time.

Data ingestion

Once again, Apache Sedona makes geospatial data ingestion remarkably easy.

				
					from pyspark.sql.functions import expr, explode, col
from sedona.spark import *
from pyspark.sql.window import Window
from pyspark.sql import functions as F
from pyspark.sql import SparkSession

config = SedonaContext.builder() .\
    config('spark.jars.packages',
           'org.apache.sedona:sedona-spark-shaded-3.3_2.12:1.7.1,'
           'org.datasyslab:geotools-wrapper:1.7.1-28.5'). \
    getOrCreate()

sedona = SedonaContext.create(config)


file_urls = {"2022": f"s3://{dataset_bucket_name}/geospatial-dataset/raster/orthophoto/soma/2022/2022_4BandImagery_SanFranciscoCA_J1191044.tif", 
              "2025": f"s3://{dataset_bucket_name}/geospatial-dataset/raster/orthophoto/soma/2025/2025_4BandImagery_SanFranciscoCA_J1191043.tif"}

df_image_2025 = sedona.read.format("binaryFile").load(file_urls["2025"])
df_image_2025 = df_image_2025.withColumn("raster", expr("RS_FromGeoTiff(content)"))

df_image_2022 = sedona.read.format("binaryFile").load(file_urls["2022"])
df_image_2022 = df_image_2022.withColumn("raster", expr("RS_FromGeoTiff(content)"))
				
			

Exploring the data

We can retrieve metadata from the 2025 image by executing the following query:

				
					
df_image_2025.createOrReplaceTempView("image_new_vw")
display(spark.sql("""SELECT RS_MetaData(raster) AS metadata, 
                  RS_NumBands(raster) AS num_bands,
                  RS_SummaryStatsAll(raster) AS summary_stat,
                  RS_BandPixelType(raster) AS band_pixel_type,
                  RS_Count(raster) AS count 
                  FROM image_new_vw"""))

				
			

Tiling for scalable image analysis

To improve scalability and enhance performance in the upcoming analysis, we first tile the image into smaller segments:

				
					
tile_w = 195
tile_h = 191

tiled_df_2025 = df_image_2025.selectExpr(
  f"RS_TileExplode(raster, {tile_w}, {tile_h})"
).withColumnRenamed("x", "tile_x").withColumnRenamed("y", "tile_y").withColumn("width", expr("RS_Width(tile)")).withColumn("height", expr("RS_height(tile)"))
window_spec = Window.orderBy("tile_x", "tile_y")
tiled_df_2025 = tiled_df_2025.withColumn("rn", F.row_number().over(window_spec)).withColumn("year", lit(2025))

tiled_df_2022 = df_image_2022.selectExpr(
  f"RS_TileExplode(raster, {tile_w}, {tile_h})"
).withColumnRenamed("x", "tile_x").withColumnRenamed("y", "tile_y").withColumn("width", expr("RS_Width(tile)")).withColumn("height", expr("RS_height(tile)"))
window_spec = Window.orderBy("tile_x", "tile_y")
tiled_df_2022 = tiled_df_2022.withColumn("rn", F.row_number().over(window_spec)).withColumn("year", lit(2022))

union_raster = tiled_df_2022.unionByName(tiled_df_2025, allowMissingColumns=False)
window_spec = Window.partitionBy("rn").orderBy(F.desc("year"))
union_raster = union_raster.withColumn("index", F.row_number().over(window_spec))
				
			

Image classification

We perform pixel-wise classification by calculating the NDVI and NDWI indices and applying a simple decision tree. The classification criteria are informed not only by the definitions of NDVI and NDWI, but also by temporal differences between the two images, which were captured at different times of the year and day. These differences are observable, for instance, in the varying shadows cast by buildings. As with any classification approach, some degree of misclassification is expected. Then the results are written to a Delta table.

				
					# Calculating NDVI using Red and NIR bands as NDVI = (NIR - Red) / (NIR + Red)
union_raster = union_raster.withColumn(
    "ndvi",
    expr(
        "RS_Divide("
        "  RS_Subtract(RS_BandAsArray(tile, 1), RS_BandAsArray(tile, 4)), "
        "  RS_Add(RS_BandAsArray(tile, 1), RS_BandAsArray(tile, 4))"
        ")"
    )
)


# Calculating NDWI using Green and NIR bands as NDWI = (Green - NIR) / (Green + NIR)
union_raster = union_raster.withColumn(
    "ndwi",
    expr(
        "RS_Divide("
        "  RS_Subtract(RS_BandAsArray(tile, 4), RS_BandAsArray(tile, 2)), "
        "  RS_Add(RS_BandAsArray(tile, 4), RS_BandAsArray(tile, 2))"
        ")"
    )
)

# Red and Green bands as arrays in new columns
union_raster = union_raster.withColumn(
    "red",
    expr(
        "RS_BandAsArray(tile, 1)"
    )
).withColumn(
    "green",
    expr(
        "RS_BandAsArray(tile, 2)"
    )
)


# Classification tree based on Red and Green bands and NDVI, NDWI
union_raster = union_raster.withColumn(
    "classification",
    F.expr("""
        transform(
            arrays_zip(ndvi, ndwi, red, green),
            x -> 
                CASE 
                    WHEN year = 2022 THEN
                        CASE 
                            WHEN x.red < 15 AND x.green < 15 THEN 4
                            WHEN x.ndvi > 0.35 AND x.ndwi < -0.35 THEN 2
                            WHEN (x.ndvi < -0.2 AND x.ndwi > 0.35) OR (x.red < 15 AND x.ndwi > 0.35) OR (x.ndwi > 0.45) THEN 3
                            WHEN x.ndvi >= -0.3 AND x.ndvi <= 0.3 AND x.ndwi >= -0.3 AND x.ndwi <= 0.3 THEN 1
                            ELSE 999
                        END
                    WHEN year = 2025 THEN
                        CASE 
                            WHEN x.red < 15 AND x.green < 15 THEN 4
                            WHEN x.ndvi > 0.3 AND x.ndwi < -0.15 THEN 2
                            WHEN x.ndvi < -0.35 AND x.ndwi > 0.55 THEN 3
                            WHEN (x.ndvi >= -0.5 AND x.ndvi <= 0.5 AND x.ndwi >= -0.5 AND x.ndwi <= 0.5) OR (x.ndvi > 0.8 AND x.ndwi > 0.3) THEN 1                           
                            ELSE 999
                        END
                    ELSE 999
                END
        )
    """)
)

# Classification array as a new band in the raster and defining no data value as 999 
classification_df = (
    union_raster
    .select("tile_x", "tile_y", "rn", "year", "index", expr("RS_MakeRaster(tile, 'I', classification) AS tile").alias("tile"))
    .select("tile_x", "tile_y", "rn", "year", "index", expr("RS_SetBandNoDataValue(tile,1, 999, false)").alias("tile"))
    .select("tile_x", "tile_y", "rn", "year", "index", expr("RS_SetBandNoDataValue(tile,1, 999, true)").alias("tile"))
)

classification_df = classification_df.withColumn("maxValue", expr("""RS_SummaryStats(tile, "max", 1, false)"""))

classification_df.withColumn("raster_binary", expr("RS_AsGeoTiff(tile)")).select("tile_x", "tile_y","rn", "year", "index", "raster_binary").write.mode("overwrite").saveAsTable("geospatial.soma.classification")
				
			

Filling missing data via interpolation

In the images above, we can observe white areas, i.e. pixels that couldn’t be classified and were left as no-data values. To address this, we can apply interpolation using the Inverse Distance Weighting (IDW) method. Apache Sedona simplifies this process: we filter out the tiles with no-data values and perform the interpolation. After that, we store the interpolated results in a Delta table.

				
					

# Separating the dataframe into two dataframes based on the 
no_interpolation_df = classification_df.filter(classification_df["maxValue"] != 999).select("tile_x", "tile_y","rn", "year", "index", "raster_binary")
interpolated_df = classification_df.filter(classification_df["maxValue"] == 999).select("tile_x", "tile_y","rn", "year", "index", expr("RS_Interpolate(tile, 2.0, 'variable', 48.0, 6.0)").alias("tile")).withColumn("raster_binary", expr("RS_AsGeoTiff(tile)")).select("tile_x", "tile_y","rn", "year", "index", "raster_binary")
union_df = interpolated_df.unionByName(no_interpolation_df, allowMissingColumns=False)
union_df.write.mode("overwrite").saveAsTable("geospatial.soma.interpolation")
union_df = spark.table("geospatial.soma.interpolation").withColumn("tile", expr("RS_FromGeoTiff(raster_binary)"))

				
			

Evaluating classification changes between the two images

It would be useful to create a band that highlights the differences between the classifications from 2022 and 2025, and add it as the third band in the image.

				
					
union_df.createOrReplaceTempView("union_df_vw")

merged_raster = sedona.sql("""
    SELECT rn, RS_Union_Aggr(tile, index) AS raster
    FROM union_df_vw
    GROUP BY rn
""")

merged_raster.createOrReplaceTempView("merged_raster_vw")
diff_raster = merged_raster.withColumn("diff_band", expr( 
        "RS_LogicalDifference("
        "RS_BandAsArray(raster, 1), RS_BandAsArray(raster, 2)"
        ")"))

result_df = diff_raster.select("rn", expr("RS_AddBandFromArray(raster, diff_band) AS raster").alias("raster")).withColumn("raster_binary", expr("RS_AsGeoTiff(raster)"))
result_df.select("rn", "raster_binary").write.mode("overwrite").saveAsTable("geospatial.soma.change_detection")
				
			

Finally, we can detect and visualize changes, such as those in vegetation, using the image below as an example:

What is next?

In our live training Databricks Geospatial in a Day at RevoData Office, we’ll delve deeper into the logic behind this code and use this example to demonstrate how to:

  • Visualize GeoTIFF files
  • Generate optimally sized tiles from large raster datasets
  • Configure clusters optimized for compute-intensive workloads, such as spatial interpolation
  • Partition Spark DataFrames by the raster column to accelerate processing
  • Use complementary tools to seamlessly merge tiles back into a single large GeoTIFF

Go ahead and grab your spot for the training using the link below, can’t wait to see you there!
https://revodata.nl/databricks-geospatial-in-a-day/

Picture of Melika Sajadian

Melika Sajadian

Senior Geospatial Consultant at RevoData, sharing with you her knowledge about Databricks Geospatial

]]>
https://revodata.nl/building-a-geospatial-time-machine/feed/ 0