Process Mining 101: Beyond Data Mining (Part 1 — What and Who)
I was invited to speak at a webinar last 3 years on Process Mining. It is a research domain introduced to me by my first Master student…
Process Mining 101: Beyond Data Mining (Part 1 — What and Who)
I was invited to speak at a webinar last 3 years on Process Mining. It is a research domain introduced to me by my first Master student. He was interested in Learning Analytics at that time, and he came across the concept of Process Mining, hence he proposed it to be his Master research topic. Ever since then, I delved deeper into the domain, and even established the Process Mining Research Cluster (PMineReC).
For the webinar, I summarized a book then newly published by Jun Matsuo (2021) called “An Introduction to Process Mining in Practice and Improvement Approach from BPM Perspective”. This article summarizes the content that I presented during the webinar.
What is Process Mining?
Process Mining can be described as a methodology of data analysis aimed at understanding business processes. It extracts event logs, or transaction data, which are stored in the system’s database, to conduct various analyses.
Imagine if the business is executed through some system, in most cases, event logs are accumulated, and these data logs are mined to analyze the process that is going on in the system. Even though business processes are the main target, process mining can be applied to a wide range of fields.
This may sound mind-boggling to those who are less familiar with the technicality of data mining and business process management. Let’s make this simpler. Every time when you use a computer system in your workplace, for example an HR system that allows you to apply for annual leave. The moment you log in to the system, submit for annual leave application, and receive notification that the application is approved, the system records these transactions in the event logs. Upon analyzing the data, process mining tool would visualize the process of the transactions from beginning to end. From the visualization, the tool would highlight the waiting time or the delay between any particular transactions.
Simply said, process mining is a data analysis method for Business Analysis. Business Analysis refers to the activity of collection and analyzing various information and data related to current business operations. The purpose of Business Analysis is to understand the mechanisms and cause-and-effect relationships of business execution, such as current work procedures and management systems. On the other hand, Process Mining mainly analyzes “business processes” among the wide range of business aspects targeted by business analysis. This is because Process Mining was originally created as an analysis method to clarify current (As-Is) business processes in the framework of Business Process Management (BPM).
In conventional business analysis, information and data related to business operations are collected mainly through reading business-related documents, such as various system specifications and manuals (document analysis), interviews and workshops with staff in charge or SME (Subject Matter Experts), time measurement with stopwatches, and observation surveys using video and other recording devices.
As mentioned above, Process Mining analyzes digital data called “event log”. It is a generic term for the operation history data, or transaction data, of systems and applications that are recorded in various IT systems used for business execution. In the case of HR system discussed above, each transaction is recorded as data with time stamps, such as date, time and up to the minutes and seconds. These records are collectively called “event logs” because they can be called “events” that occur in system operations.
Since the volume of event log data extracted from such systems is sometimes tens to hundreds of gigabytes, it can be said that process mining is a kind of big data analysis.
Process mining can be considered as a kind of “data mining”. It is a bridge between data science and process science. If you are familiar with data mining, it is a general-purpose analysis method that handles all kinds of big data from various fields. Process mining differs in that it literally focuses on processes, not discrete data per se. The essence that differentiate data mining and process mining is “time”. With the element of time, process mining could analyze any data to visualize the significant events.

As As shared by Ernesto Damiani, AI and Intelligent Systems Institute, Khalifa University, UAE
Time series analysis that incorporates time stamps can be used to measure business efficiency. As an example, how many days, hours and minutes the processing time of an individual activity (an event), such as an annual leave application, or an approval for the annual leave. Conventional business analysis would be very time-consuming and costly, as the number of surveys is limited to a few dozen to a few hundred samples at most. Process mining, however, extracts hundreds of thousands to millions of event logs from IT systems, so it is practically a survey of the entire system, and it enables us to grasp the current situation very close to reality.
We extract event logs from business systems and analyze them using three basic approaches:
- Process Discovery: It automatically draws the current process (As-Is) as a flowchart. The current process visualized as a flowchart is analyzed from various perspectives to identify inefficient processes, bottlenecks, etc. It is the fundamental analysis of process mining.
- Conformance Checking: It compares and analyzes the current process (As-Is process) grasped through process discovery with the ideal process (To-Be process), and identifies deviations and violations in the current process when the ideal process is considered as positive.
- Performance Enhancement: This is an effort to improve the process by correcting problem areas identified through process discovery and conformance testing.
Who is behind Process Mining?
We should always give credit to the founder. Without the idea that he brought forth, we would have been struggling with mining data for process enhancement. The name to remember is Wil van der Aalst.
A Dutch professor at RWTH Archen University, Prof van der Aalst, started studying workflow and workflow management at the Eindhoven University of Technology (TU/e) in the Netherlands, in 1990s. Realizing that the existing methods (at that time, interviews and workshops) for understanding business processes could only draw incomplete process models based on subjective and fragmented information. Prof van der Aalst’s main areas of expertise are information systems (IS), workflow management, and process mining. Process mining is a young technology, born in 1999.
In the 1990s, business systems such as SAP’s ERP operations in various departments of companies were being conducted on IT systems. Prof van der Aalst used the term “process mining” for the first time in his research proposal in 1998, and began his research on this in 1999. Since the early 2000s, academic research has been actively conducted at universities in Europe, especially at TU/e, where Aalst was a full professor then.
On the receiving end, Anne Rozinat and Christin W. Gunther developed a process mining tool called “Disco” (stands for Discovery) as they founded Fluxicon in 2010 after obtaining their PhDs at TU/e. With the growth of process mining tools since 2009, Rozinat and Gunther are among those who contributed to increasing the awareness and understanding of process mining in Europe. Honestly speaking, if you want to explore process mining, Fluxicon Disco is a good tool to start with.
A bit of chronological history here. The ProM tool, a software for process mining, was developed by Willem Groenewegen, Jan Willem van der Aalst, and their colleagues at the Eindhoven University of Technology (TU/e), and the tool’s development began around 2003. In 2011, van der Aalst published “Process Mining: Data Science in Action” under Springer publication, and it has a few editions since. In 2014, van der Aalst developed and launched a MOOC (e-learning course) on the online learning platform Coursera, with the same title as the book. Tens of thousands of people around the world has taken up this e-learning course. In 2019, the International Conference on Process Mining was held for the first time in Archen, Germany. Due to the pandemic, the conference was held online in 2020.
Geographically, process mining has gained interests across the globe. Among the early adopters are Australia, Japan, the United States of America (USA), and South Korea. In Australia, an active effort led by researchers at the University of Melbourne, where an open-sourced process mining tool called “Apromore” was developed. Due to the exposure and high interest at the Process Mining Conference 2019, the two tools with largest presence in the market since, Celonis and myInvenio, were born. ABBYY Timeline and Signavio are another two that are actively working to expand their presence in Japan. Generally, in Asia, more exposure is needed. A group of researchers who studied under van der Aalst is building up a track record of introducing process mining to companies in South Korea.
It has been a long write-up on the What and Who in process mining. I reserve the rest of the details, the Why and How, in my next article.
Stay tuned!
메타데이터
- post_id
- cc48dd797e2c
- slug
- process-mining-101-beyond-data-mining-part-1-what-and-who-cc48dd797e2c
- url
- https://medium.com/@sha905/process-mining-101-beyond-data-mining-part-1-what-and-who-cc48dd797e2c
- canonical_url
- https://medium.com/@sha905/process-mining-101-beyond-data-mining-part-1-what-and-who-cc48dd797e2c
- author_url
- https://medium.com/@sha905
- status
- ok
- fetched_at
- 2026-06-11 21:11:36