Analyzing the Impact of Increasing Daily Wages of Tea Plantation Labourers to LKR 2000
The tea industry is of immense importance to Sri Lanka, as it employs thousands of workers and generates significant revenue for the…
Analyzing the Impact of Increasing Daily Wages of Tea Plantation Labourers to LKR 2000

The tea industry is of immense importance to Sri Lanka, as it employs thousands of workers and generates significant revenue for the country. In my previous article, titled ‘A Data-Driven Strategy for Sri Lanka to Achieve 10 Billion USD in Annual Exports,’ I identified the tea industry as the second most suitable sector for Sri Lanka to achieve this target, with the potential to generate 113.33 million USD in export income per month.
Today, we are focusing on data collection and preparation to analyze the impact of increasing the daily wages of tea plantation laborers to LKR 2000. In this initial phase, it is crucial to ensure that this wage increase does not compromise the industry’s ability to generate export income. Therefore, conducting this analysis is both relevant and worthwhile.
This section focuses on building the necessary dataset for the analysis. It details the use of official data sources, describes the methods employed to collect the information, and explains how the data was cleaned and prepared. The resulting robust dataset will then be utilized in the next stage of the section Data Analysis and Interpretation, where a comprehensive analytical study will be conducted to thoroughly examine the consequences of the proposed LKR 2000 wage hike.
Data sources used
The Department of Census and Statistics data
- The Department of Census and Statistics, Sri Lanka, provided the “Cost of Production of Tea per Kilogram, 2019/20–2023/24” table (available at: https://www.statistics.gov.lk/Agriculture/StaticalInformation/COP_tea ). This table contains data on the production cost per kilogram of tea for the financial years 2019/20 to 2023/24.
- The ‘Production and Cost of Production (COP) of Tea, Rubber, and Coconut: 1994–2009’ webpage data was also extracted from https://www.statistics.gov.lk/Agriculture/StaticalInformation/National_TeaRubberCoconut . It provides the cost of production (COP) per kilogram in Rupees (Rs), as well as the extent (in hectares) and production volume (in metric tons, MT) for the years spanning from 1994 to 2009.
- The statistical pocket book 2024, available at https://www.statistics.gov.lk/Publication/PocketBook2024 , provides the financial year-wise (2012/13–2022/23) production cost per kilogram of tea. This data was utilized to fill in the missing production cost values for tea from the financial years 2012/13 to 2019/20, as mentioned in the first source provided by the Department of Census and Statistics.
Department of Labour Data
- The ‘Labour Statistics Sri Lanka 2018’ (available at https://labourdept.gov.lk/downloads/stat/10.pdf ) and the ‘Annual Labour Statistics Report — 2021’ (available at https://labourdept.gov.lk/downloads/stat/7.pdf ) are used to gather data on the Minimum Daily Rate of Wages (in Rupees) and the Total Wage (per day) for the year span of 2010 to 2021.
Central Bank Reports Data
- The ‘Sri Lanka Socio Economic 2018’ report data (available at https://www.cbsl.gov.lk/sites/default/files/cbslweb_documents/statistics/SriLanka%20Socio_Economic_Data_2018_e.pdf ) and the ‘Sri Lanka Socio Economic 2025’ report data (available at https://www.cbsl.gov.lk/sites/default/files/cbslweb_documents/publications/otherpub/publication_sri_lanka_socio_economic_data_folder_2025_e.pdf ) are used to obtain the selling price (Selling Price Colombo Auction (Gross) per kilogram). The data spans from 2012 to 2024.
Research Paper Data
- The research paper titled ‘FACTORS INFLUENCING LABOUR TURNOVER IN THE TEA PLANTATION SECTOR IN BADULLA REGION IN SRI LANKA: A CASE STUDY ON CULLEN TEA ESTATE, BADULLA’ (available at https://ours.ou.ac.lk/wp-content/uploads/2022/12/ID-129-FACTORS-INFLUENCING-LABOUR-TURNOVER-IN-THE-TEA-PLANTATION-SECTOR-IN-BADULLA-REGION-IN-SRI-LANKA-A-CASE-STUDY-ON-CULLEN-TEA-ESTATE-BADULLA-.pdf ) is used to obtain tea labour force data. It contained the total number of laborers for the years 1988, 1990, 2011, and 2018.
- The Anker Research Institute’s ‘2021/2022/2023/2024 Living Wage Update: Estate Sector, Sri Lanka’ reports (available at https://www.globallivingwage.org/living-wage-benchmarks/living-wage-estimate-for-sri-lanka/ ) were used to obtain the ‘Living Wage per Month’ and the ‘Cost of a Decent Standard of Living for a Family per Month’ for estate sector workers. The data covers the year span of 2015 to 2024.
Other Open Data sets
- To gain an understanding of inflation over the years, the Consumer Price Index (CPI) is a crucial metric. This data was obtained from the Macrotrends website, available at https://www.macrotrends.net/global-metrics/countries/lka/sri-lanka/inflation-rate-cpi , and it includes information spanning from 1988 to 2024.
- To gain an understanding of the unemployment rate in Sri Lanka, data was obtained from the Macrotrends website, available at https://www.macrotrends.net/datasets/global-metrics/countries/lka/sri-lanka/unemployment-rate . The data spans from 1991 to 2024.
- The Gazette of the Democratic Socialist Republic of Sri Lanka was also used to obtain the latest updates on tea labor wages, available at https://documents.gov.lk/view/extra-gazettes/2024/8/2397-27_E.pdf .
- The Wikipedia dataset on tea hectares, available at https://en.wikipedia.org/wiki/Tea_production_in_Sri_Lanka , the Tea Exporters Association Sri Lanka (TEA) dataset on tea hectares, available at https://teasrilanka.org/statistics , and the annual tea production data from 1970 to 2017 in the chart available at https://www.ceicdata.com/en/sri-lanka/production-by-commodity/production-annual-tea , were used to obtain the total hectares of tea cultivated. This data was then combined with the previously mentioned ‘Production and Cost of Production (COP) of Tea, Rubber, and Coconut: 1994–2009’ dataset to complete the overall dataset.
- The report ‘Ceylon Tea Industry and Regional Plantation Companies’ available at https://www.historyofceylontea.com/pdf-load/Articles/rr-ceylon-tea-rpcs-2020-min.pdf provided data on ‘total wages per day’ and ‘Selling Price/Kg’ for the years 1988–2019 and ‘total laborers’ for several years within the same period. This data was used to fill in the missing fields in my dataset. Additionally, the report provided ‘profit per kilogram’ data, which was also included in my dataset.
Data Collection methods
For this study the data were collected from secondary sources and from the above-mentioned sources. Each dataset was carefully identified, downloaded, and compiled to ensure consistency, reliability, and relevance to the topic — the impact of increasing the daily wage of tea plantation labourers to LKR 2,000.
- The Department of Census and Statistics website provides data on the cost of production of tea per kilogram. This data was copied and pasted into an Excel sheet, with the financial year used as the index column. The remaining missing data for the years 2012/13 to 2019/20, which included the production of tea, were obtained from the Statistical Pocket Book 2024. This ensured continuity of the cost-of-production variable based on the financial year. On the other hand, the same cost of production per kilogram of tea was available on a calendar year basis for a larger number of years on the ‘Production and Cost of Production (COP) of Tea, Rubber, and Coconut: 1994–2009’ webpage. This data was also extracted and saved to the same Excel file but under a new index called ‘Year.’ It was recorded under a new variable, ‘Production Cost per kg (Calendar Year).’
- Labour wage data were extracted from the Department of Labour’s official annual reports (2018 and 2021), which were available in PDF format. Relevant tables listing the minimum daily wage and total daily wage by year were converted into tabular data and manually cross-checked for accuracy. The data from both 2018 and 2021, which shared similar column names, were merged and entered into the same Excel sheet under a unified “calendar year” column, using the year as the index. Additionally, the ‘Ceylon Tea Industry and Regional Plantation Companies’ report provided the same data in calendar year format, which was also integrated into the dataset. However, in cases of overlap or discrepancies, priority was given to the data from the Department of Labour over the ‘Ceylon Tea Industry and Regional Plantation Companies’ report.
- The ‘Ceylon Tea Industry and Regional Plantation Companies’ PDF provided the ‘profit per kilo’ data, which was also compiled into my dataset using the calendar year column as the base.
- Selling price data were collected from the CBSL Sri Lanka Socio-Economic Data reports (2018 and 2025 editions), specifically using the “Selling Price Colombo Auction (Gross) per kg” table. These values were recorded year-wise and matched to corresponding production years in the dataset. Additionally, the data was further enriched using information from the ‘Ceylon Tea Industry and Regional Plantation Companies’ report.
- Macroeconomic indicators, such as the inflation rate (CPI) and unemployment rate, were obtained from the Macrotrends database, which provides historical datasets spanning from 1988 to 2024. These datasets were downloaded in CSV format and integrated into the main dataset using the ‘calendar year’ column to facilitate comparative economic analysis.
- Additionally, the Government Gazette (Extraordinary Gazette №2397/27, 2024) was consulted to capture the latest officially declared daily wage for tea estate workers. The Gazette was accessed online and served as the most up-to-date legal reference for the study. The relevant value was manually entered into the dataset to ensure accuracy and alignment with official records.
- Data on the labour force size and turnover were gathered from the research paper “Factors Influencing Labour Turnover in the Tea Plantation Sector in Badulla Region” (Open University of Sri Lanka). Numerical data from the case study, which included labour counts for the years 1988, 1990, 2011, and 2018, were manually entered into the dataset and referenced using the ‘calendar year’ column. Additionally, the ‘Ceylon Tea Industry and Regional Plantation Companies’ report provided similar data on a yearly basis. However, in cases of overlap or discrepancies, priority was given to the data from the research paper.
- Living wage and household cost-of-living information were collected from the Anker Research Institute’s Living Wage Updates for the Estate Sector (2015–2024) and it was manually entered to the source excel referencing the ‘calendar year’ column.
The data I collected, as mentioned above, were assigned priorities when two different sources provided data for the same year and attribute. The first priority was given to data obtained from the Department of Census and Statistics, the Department of Labour, Central Bank reports, and the Gazette. Medium priority was assigned to data from research papers and the Anker Research Institute. The lowest priority was given to data obtained from open datasets, such as Wikipedia, and reports like the Ceylon Tea Industry and Regional Plantation Companies report.
Considering all the above sources and the collection and compilation methods described, the following is the dataset I created before proceeding with the preprocessing.

Table 1 — Dataset before Preprocessing
Data Preprocessing
As illustrated in Table 1, after compiling data from various sources using different data collection methods, the dataset contains numerous issues such as missing values and data duplications. Therefore, it is essential to clean the dataset and address these issues to ensure its suitability for conducting data analysis and interpretation in next section.
Missing data can be addressed primarily through two approaches: deletion and imputation. However, given the small size of this dataset, deletion (including methods such as listwise deletion or pairwise deletion) is not an ideal option, as it could lead to a significant loss of valuable information. Therefore, imputation is the most suitable approach in this case.
Before proceeding with imputation, it is essential to identify the nature of the missingness in the dataset — whether it falls under MCAR (Missing Completely At Random), MAR (Missing At Random), or MNAR (Missing Not At Random). I made every effort to recover the missing data by reviewing various data sources and conducting extensive internet research, but unfortunately, this attempt was unsuccessful. Since the missingness appears unrelated to any other variables in the dataset and seems to have occurred purely due to lost files, it can be categorized as MCAR (Missing Completely At Random). The good news is that in such scenarios, the analysis remains unbiased, and appropriate imputation methods can be applied to handle the missing data effectively.
1. Redundancy Removal
- As illustrated in Figure 1, the dataset contains the same data represented under two different names: the “calendar year” and “financial year” columns. This is an example of data redundancy (or data duplication), and to address this, we can delete one of the columns. In this case, I have chosen to delete the “financial year” column.
- Similarly, the dataset includes the production cost per kilogram in both “production cost/kg (Financial Year)” and “production cost/kg (Calendar Year)” formats. Here too, we can eliminate one of the columns. However, before deleting one, it is necessary to first combine the data from both formats into a single column
- In the dataset, the ‘Minimum Daily Rate of Wages (Rs.)’ column and the ‘Total Wage (per day)’ column contain similar data. Therefore, we can retain one and remove the other. When making this decision, it is important to consider the objective of our task, which is to analyze the impact of increasing the daily wages of tea plantation laborers to LKR 2000. Since the focus is on the ‘total wage’ rather than the ‘minimum wage,’ and the compiled dataset contains complete data for the ‘Total Wage (per day)’ attribute, it is more appropriate to remove the ‘Minimum Daily Rate of Wages (Rs.)’ column to address data redundancy. Thus, I have chosen to keep the ‘Total Wage (per day)’ column.
- The dataset contains two attributes, ‘Living Wage/month’ and ‘Cost of Decent Standard of Living for a Family/Month,’ which essentially convey the same meaning. To avoid data duplication, one of them can be removed. In this case, I have chosen to remove the ‘Cost of Decent Standard of Living for a Family/Month’ attribute because the term ‘decent standard’ lacks a clear and specific context in this scenario. Therefore, the ‘Living Wage/month’ attribute will be retained.
2. Imputation
Here we are dealing with time series data and no categorical data are exist in here at the moment. Keeping that in the mind, we need to go through each attribute/ feature to identify the optimum imputation method.
2.1 ‘production cost/ kg’

Table 2
The dataset contains missing values scattered throughout. To examine whether there are any trends or seasonal/cyclic patterns, a time series plot was created, as shown in Figure 1.

Figure 1 — Production cost /Kg by Year
As illustrated in Figure 1, the plot reveals a clear trend but no evidence of seasonality. Based on this observation, linear interpolation could be considered as a potential method for handling missing values. However, there is a concern: the data appears to follow an exponential pattern rather than a linear one. Assuming linearity might lead to impractical results, such as negative values for the years 1988–1990. Therefore, the most appropriate approach is to fit a polynomial or exponential regression model to the dataset and use the resulting equation to impute the missing values. While mean, mode, or median imputation methods could technically be applied, they are not practical in this context, as they would likely produce unrealistic values for the earlier years.

Figure 2 — Production cost/Kg by year after fitting exponential regression line
According to Figure 2, the data fits the equation y = 1E-77e0.0909x with an R2 vale of 98%. This high R² value indicates that the equation explains a significant portion of the variance in the data. Therefore, the best course of action is to impute the missing values using this exponential regression equation.
2.2 ‘Extent (1) (Hectares)’
Before preprocessing, the annual breakdown of tea extent in Sri Lanka (in hectares) is presented in Table 3. The dataset shows repeated values for several consecutive years. For example, the value 188,970 is repeated for the years 1994–2001, and 212,715 is repeated for the years 2002–2007.

Table 3
To determine whether there are any trends or cyclic patterns in the dataset, a plot of tea extent by calendar year was generated, as shown in Figure 3.

Figure 3 — Tea Extent by year
According to Figure 3, there is no clear trend or seasonality in the data. For such scenarios, various imputation methods can be applied, including mean, median, mode, or random sample imputation. However, in this case, I have chosen to use the Last Observation Carried Forward (LOCF) method. This approach is suitable because the original dataset already exhibits repeated values over multiple years, and there is a slight upward trend.
Based on this method:
- The value 200,001 will be imputed for all missing years from 2019 onwards.
- The value 188,970 will be imputed for the missing years from 1988 to 1993.
2.3 ‘Production(1) MT’
The total tea production was represented in this column, and after compiling the data from the various sources mentioned above, the variable is as shown in Table 4, prior to any preprocessing.

Table 4
Here, only one value is missing. By examining the behavior of the attribute, it appears that mean, median, or mode imputation could be suitable. However, before making a decision, it is crucial to analyze the behavior of this time series data. A scatterplot of the data is provided in Figure 4 below.

Figure 4 — Tea Production by Year
According to Figure 4, there is no clear trend or cyclic/seasonal pattern evident in this attribute. Therefore, I have chosen to use the mean to replace the missing value by calculating the average of the values immediately before and after the missing point (Linear Interpolation), as this is a time series dataset.
The value to be entered into the missing cell is calculated as follows:
- (221,836+242,210)/2 = 232,023
2.4 ‘Total labors’
Prior to the preprocessing steps and after compiling the data from various sources, the ‘Total Labor’ involved in the tea plucking process is listed in Table 5 below. Here, we can observe missing values for several years, and it is necessary to carefully impute appropriate values for this column.

Table 5
Since this is time series data, it is essential to understand the behavior of this attribute over the years. The behavior of the data is illustrated in Figure 5 below.

Figure 5 — Tea Labor by Year
Here, we can observe a clear downward trend with no seasonal or cyclic components present. Therefore, while linear interpolation could be a viable option, considering the locations of the missing values, filling the missing values using a linear regression line would be more practical and suitable in this case.
I have fitted a linear regression line for this attribute, achieving an R² value of 86.65%. The resulting regression line is illustrated in Figure 6 below.

Figure 6 — Linear Regression on tea labors
Therefore, using the fitted regression equation, we can calculate and impute the missing values for the missing data. The regression line was fitted based on the summary output provided below, which shows a p-value of less than 0.05. This result leads to the rejection of the null hypothesis, where:
- H₀ (null hypothesis): There is no relationship between ‘total labor’ and ‘calendar year.’
- H₁ (alternative hypothesis): There is some relationship between ‘total labor’ and ‘calendar year.’

The significant p-value indicates that there is sufficient evidence to support the existence of a relationship between the two variables, justifying the use of the regression equation for imputation.
The regression equation for this case is:
y = -9005.54297958155x+18299353.9203428
By substituting the missing ‘calendar years’ (x-values) into this equation, we can calculate the corresponding imputed values (y-values) for the missing data points.
2.5 ‘unemployment rate’
After compiling data from various sources and before proceeding with data preprocessing, the ‘unemployment rate’ column for Sri Lanka is presented in Table 6. In this dataset, only the initial set of values is missing.

Table 6
Since this is time series data, it is essential to examine its behavior, as was done for previous attributes. A scatter plot of the unemployment rate is provided in Figure 7 below.

Figure 7 — Unemployment rate by Year
According to Figure 7, there is a clear downward trend in the unemployment rate, with no evidence of seasonal or cyclic patterns in this attribute. Based on this observation, the Next Observation Carried Backward (NOCB) method is an appropriate approach for handling the missing values, considering the behavior of this attribute. This means that the missing values can be replaced with the next available observation, which is 14.66%.
2.6 ‘profit per kilo’
The profit from 1 kg of tea is represented in this variable, and after compiling data from various sources, it appears as shown in Table 7. The missing data is located in the latter part of the attribute.

Table 7
To understand the behavior of this attribute, it is necessary to plot its behavior over time. This is illustrated in Figure 8 below.

Figure 8- Profit per kilo by Year
According to Figure 8, there is no clear trend, seasonal pattern, or cyclic pattern evident in this attribute. Therefore, several options are available for handling these missing values, such as mean, median, mode, or random sample imputation. Considering the nature of this attribute, I have chosen to replace the missing values with the mean.
After calculating the mean, the value obtained was -5.47, and this value was used to fill all the missing entries.
2.7 ‘Living Wage/month’
The living wage for a month for an estate worker is represented in this column in the dataset and is shown in Table 8 below, prior to data preprocessing. The column contains a significant number of missing values. Typically, one approach to handling such cases is to delete the column if possible (e.g., through pairwise deletion or dropping the variable). However, in this case, the attribute is highly important because we are analyzing the impact of increasing daily wages of tea plantation laborers to LKR 2000, as well as the potential economic, social, and industrial impacts, as required in BA 4005 Task 5. Therefore, it is essential to fill the missing values using an appropriate method.

Table 8
To determine the best approach, I visualized this variable against the calendar year to understand its behavior. The results are shown in Figure 9 below.

Figure 9 — Living Wage/month by Year
According to Figure 9, the data exhibits an upward trend with no seasonal or cyclic components. While simple linear regression might initially seem like a viable option, it is not suitable here because it could result in negative values for the earlier years, given the distribution of data points in Figure 9. To address this, I attempted to fit an exponential regression equation, as living wages typically exhibit exponential growth patterns.
The fitted plot is shown in Figure 10 below.

Figure 10 — Exponential Regression curve on Living Wage/month
The exponential curve provides an R² value of 83%, indicating a good fit. The suggested equation is:
y = 51081e^0.0967x*
This equation serves as a reliable estimate for the missing values. Using this equation, the missing values were calculated and filled accordingly.
2.8 ‘Total wage(per day)’,’ Selling Price Colombo Auction (Gross) /Kg’ and ‘cpi’
For the ‘Total wage(per day)’,’ Selling Price Colombo Auction (Gross) /Kg’ and ‘cpi’ variables, all the data are collected and no missing values are available, see the Table 9.
Just a rounding is required for the consistency of the data for these 3 attributes.

Table 9
After filling the missing values using various techniques, the dataset now appears as shown in Table 10 below.

Table 10 — Dataset after handling the missing values
3. Treating Outliers
The next step is to check for outliers in the dataset. There are several methods to identify outliers, such as the Interquartile Range (IQR) method, where values outside 1.5 times the IQR are flagged as potential outliers, and the Z-score method, where data points with a Z-score greater than +3 or less than -3 are considered potential outliers. When an outlier is flagged, it is important to investigate the reason behind it. The outlier could be due to an error in data collection or entry, or it could represent an actual value that reflects a special case. For example, if we are analyzing batting averages in cricket, Don Bradman’s average might be flagged as an outlier. However, in such cases, we cannot treat it as an outlier because it is a legitimate value that should be retained for accurate analysis.
For this analysis, I will use the IQR method to flag values outside 1.5 times the IQR as potential outliers. These flagged values will then be verified and treated accordingly. Below are the box-and-whisker plots I constructed for each attribute to identify and analyze potential outliers.

Figure 11 — Box Plot for Production Cost/Kg

Figure 12 — Boxplot for Extent (Hectares)

Figure 13 — Boxplot for Production (MT)

Figure 14 — Boxplot for Total Wage(per day)

Figure 15 — Boxplot for Selling Price

Figure 16 — Boxplot for Total labors

Figure 17 — Boxplot for Unemployment Rate

Figure 18 — Boxplot for Living Wage/month

Figure 19 — Boxplot for cpi

Figure 20 — Boxplot for profit per kilo
According to Figures 12–20, outliers have been flagged in the attributes ‘Total Wage (per day)’, ‘Selling Price Colombo Auction (Gross) /Kg’, ‘Living Wage/month’, ‘CPI’, and ‘Profit per Kilo’.
To better identify the specific values flagged as outliers, I created the following IQR tables (Tables 11 and 12) to pinpoint their locations.

Table 11 — IQR

Table 12 — Identification of Outliers
Based on Tables 11 and 12, the following values were identified as potential outliers:
-
Total Wage (per day) in 2024 — However, this is not an actual outlier. It was officially gazetted in 2024, so we can safely ignore this.
-
Selling Price Colombo Auction (Gross) /Kg in 2022–2024 — These are also not actual outliers, as they are based on government-reported official data sources. Therefore, these values should remain unchanged.
-
Living Wage/month in 2023 and 2024 — These values are supported by official data sources and are not considered outliers. They should be retained in the dataset.
-
CPI in 1990, 2008, and 2022 — These values are sourced from credible official World Bank data. As such, they cannot be challenged and must be kept in the dataset.
-
Profit per Kilo in 2011, 2015–2017, and 2019 — These values are obtained from credible sources and must also be retained in the dataset.
Since all flagged values are either verified by credible sources or represent accurate real-world data, none of them qualify as true outliers that require treatment.
Therefore, the final cleaned dataset is presented in Table 13 below.

Table 13 — Final cleaned Dataset
Final Dataset Description
Description of each attribute
The final dataset (above Table 13) contains 37 records, covering the time period from 1988 to 2024 on a yearly basis. It includes 11 attributes, which are as follows’
-
Calendar Year — This column represents the year of each record and serves as the index column of the dataset, covering the years 1988 to 2024, inclusive.
-
Production Cost per kg — This indicates the production cost of 1 kg of tea in LKR, including labor costs, manufacturing costs, marketing costs, transport costs, and other expenses incurred during production.
-
Extent (1) (Hectares) — This shows the total area, in hectares, of tea planted in Sri Lanka.
-
Production (1) (MT) — This represents the total production of tea, measured in metric tons.
-
Total Wage (per day) — This reflects the total wage that a tea worker can earn per day. It is calculated by adding the ‘Daily Attendance Incentive (Rs.)’, ‘Daily Price Share Supplement (Rs.)’, ‘Budgetary Relief Allowance (Rs.)’, or ‘Productivity-Based Incentive’ to the ‘Minimum Daily Rate of Wages’.
-
Selling Price at Colombo Auction (Gross) per kg — This indicates the gross selling price of 1 kg of tea at the Colombo tea auction.
-
Total Laborers — This shows the total number of tea plantation laborers.
-
Unemployment Rate — This represents the unemployment rate in Sri Lanka. It was included because tea plantation workers are more likely to face unemployment if they leave their jobs.
-
Living Wage per Month — This represents the bare minimum labor wage for a tea plantation worker. It is calculated by adding EPF, Kovil Fund, Union Dues, and PAYE (if applicable) to the Net Living Wage.
-
CPI — This is the ‘Consumer Price Index’ in Sri Lanka, representing inflation. It reflects the annual percentage change in the cost to the average consumer of acquiring a basket of goods and services, which may be fixed or updated at specified intervals, such as yearly.
-
Profit per Kilo — This shows the profit or loss a tea company earns from selling 1 kg of tea.
Summary statistics
The summary statistics for the attributes are presented in Table 14.

Table 14 — Summary Statistics
The dataset has now been fully preprocessed and is ready for the next section : Analyzing the Impact of Increasing Daily Wages of Tea Plantation Labourers to LKR 2000 — Data Analysis and Interpretation.
In this Section, I will focus on the data analysis and interpretation. The analysis will cover the potential economic, industrial, and social impacts of the wage increase. For this purpose, I will use descriptive, diagnostic, and predictive analysis to examine the economic and industrial impacts. Additionally, for the social impact, I will employ a simulation model (NetLogo) to study labor behavior in a controlled environment.
The dataset I am using for this analysis, after preprocessing (as described in the previous task), is shown in Table 1.

Table 1 — Data Set Used for Analysis
Economic and Industrial Impact Analysis
Here, we will analyze the economic and industrial impact of increasing the daily wages of tea plantation laborers to LKR 2000, with respect to production costs, total tea production, selling price of tea, and profitability. Additionally, we will examine how the extent of tea cultivation in Sri Lanka has behaved over the past years. For this purpose, we will begin with a descriptive analysis, followed by a diagnostic analysis. These analyses will then be used to explore the social impact of the wage increase decision, which will be covered in the next section.
First, let us examine the big picture of the Tea Plantation Companies’ Key Performance Indicators (KPIs) during the period from 1988 to 2024. This is illustrated in Figure 1 below.

Figure 1 — Overall Big Picture
According to the data, the average production cost per kilogram during the period of 1988–2024 was LKR 279.51, with an average selling price of LKR 342.43 per kilogram of tea. However, it resulted in an average loss of LKR 5.47 per kilogram during this period, which is an interesting revelation. The year-over-year (YoY) graph of profit, selling price, and production cost per kilogram further highlights that profits turned negative in the later years, especially after 2010. In 2019, there was a significant loss of LKR 106.37 per kilogram. Additionally, the data shows that even after 2022, despite production costs being significantly lower compared to the selling price, the industry still incurred a loss of LKR 5.47 per kilogram. This raises an intriguing question: Can tea companies withstand an increase in the daily wages of tea plantation laborers to LKR 2000? With this in mind, let us proceed to examine other details.
The below Figure 2 illustrates the Total Tea Production in Metric Tons per one hectare of land extent.

Figure 2 — YoY Production (MT) per Hectare
As you can see here, the average tea plucked (in metric tons) from one hectare has been gradually increasing. This indicates that, with respect to tea production, the industry has been able to maximize output from each hectare during the period of 1988–2024, which is a positive sign. This practice allows them to save a considerable amount by optimizing land use.
Now, let’s dive deeper to examine how these variables — specifically, production cost, total tea production (in metric tons), selling price, and profit — are correlated with the variable ‘Total Wage per Day’ for a tea plantation laborer. This is crucial because tea companies are ultimately responsible for paying these workers, and any analysis must ensure that the question of whether tea companies can withstand an increase in daily wages to LKR 2000 is not compromised.
Since we already have an understanding of the behavior of each variable (discussed under BA4005 Task 4, where the behavior of each variable was plotted against the year during the data preprocessing phase), there is no need to repeat that here. Moving forward, I will focus on analyzing the correlation between the ‘Total Wage (per day)’ variable and other variables to gain insights into the economic and industrial impact of the decision to increase wages.
Total wage(per day)’ and ‘production cost/ kg’
In order to examine the relationship between these two variables, we perform a regression analysis using simple linear regression. Let us analyze how these two variables are associated, as illustrated in Figure 3 below.

Figure 3 — Production cost/kg by Total Wage per Day
From this, we can observe that there is a linear relationship between the two variables. Based on this, we can formulate the following hypotheses:
- H0: There is no relationship between Total Wage (per day) and production cost per kg.
- H1: There is some relationship between Total Wage (per day) and production cost per kg.
The Excel output is as follows:

Table 2 — Summary Output for Total Wage (per day) vs Production cost/ Kg
According to Table 2, the R² value is 0.88, indicating that 88% of the variation in production cost per kg can be explained by Total Wage (per day). Additionally, the p-value is less than 0.05, allowing us to reject H0 and conclude that there is a significant relationship between production cost per kg and Total Wage (per day).
Furthermore, we can examine the residual plots below for model diagnostics.

Figure 4 — Model Diagnosis for Production Cost/Kg vs Total Wage (per day)
As shown in Figure 4, we observe that:
- The average of the residuals is 0.
- The spread of the residuals is even.
- The residuals do not follow a specific pattern or trend.
- The errors are normally distributed, as confirmed by the normality probability plot.
Based on these findings, the final Simple Linear Regression (SLR) equation is: ŷ (Production Cost per Kg) = 29.8 + 0.64 × Total Wage (Per Day)
‘Total wage(per day)’ and ‘Production(MT)’
In order to determine the relationship between these two variables, we first plot them on a scatterplot, as shown in Figure 5 below.

Figure 5 — Production (MT) by Total Wage per Day
According to the scatterplot, there appears to be no clear linear relationship between these two variables. However, let us proceed to examine the summary outputs to formally test the hypotheses:
- H0: There is no relationship between Total Wage (per day) and production (MT).
- H1: There is some relationship between Total Wage (per day) and production (MT).
The Excel output is as follows:

Table 3 — Summary Output Total Wage (per day) vs Production (MT)
According to Table 3, the R² value is 0.05, indicating that only 5% of the variation in tea production (MT) can be explained by Total Wage (per day). This is a very low value, and since the p-value is greater than 0.05, we do not have sufficient evidence to reject the null hypothesis (H₀) at the 95% confidence interval. Therefore, it can be concluded that there is no significant relationship between Total Wage (per day) and tea production (MT).
As a result, there is no need to check the residuals, and it would be unnecessary to derive a Simple Linear Regression (SLR) equation for these variables.
‘Total wage(per day)’ and ‘Selling Price/Kg’
To determine the relationship between Total Wage (per day) and Selling Price per Kg, we similarly need to plot a scatterplot for these two variables. The scatterplot, shown in Figure 6 below, indicates that there is a linear relationship between the two variables.

Figure 6 — Selling Price/Kg) by Total Wage per Day
From Figure 6, we can observe a linear relationship between the two variables. Therefore, we should proceed to test the hypotheses using regression analysis. The hypotheses of interest are as follows:
- H₀: There is no relationship between Total Wage (per day) and Selling Price per Kg.
- H₁: There is some relationship between Total Wage (per day) and Selling Price per Kg.

Table 4 — Summary Output Total Wage (per day) vs Selling Price/Kg
The Excel output is as follows:
According to Table 4, the R² value is 0.886, which indicates that 88.6% of the variation in the selling price per kg can be explained by the Total Wage (per day). Additionally, since the p-value is less than 0.05, we can reject H₀ and conclude that there is a significant relationship between the selling price per kg and Total Wage (per day).
Furthermore, we can examine the model diagnostics, as shown in Figure 7.

Figure 7- Model Diagnosis for Selling Price per Kg vs Total Wage (per day)
Here, we can observe the following:
- The average of the residuals is 0.
- The spread of the residuals is even.
- The residuals do not follow a specific pattern or trend.
- The errors are normally distributed, as confirmed by the normality probability plot.
Based on these findings, the final Simple Linear Regression (SLR) equation is: ŷ (Selling Price per Kg) = 24.6 + 0.82 × Total Wage (Per Day)
‘Total wage(per day)’ and ‘Profit Per Kg’
Here, we also need to plot the behavior of the two variables, Total Wage (per day) and Profit Per Kg, on a scatterplot. This is shown in Figure 8 below.

Figure 8 — Profit per kilo by Total Wage per Day
According to the above figure, it appears that there is no relationship between these two variables, as the data points are randomly scattered around the x-axis. However, let us test the hypotheses below for further clarification:
- H₀: There is no relationship between Total Wage (per day) and Profit per Kg.
- H₁: There is some relationship between Total Wage (per day) and Profit per Kg.
The Excel output is as follows:

Table 5 — Summary Output Total Wage (per day) vs Profit per kg
According to this table, the R² value is 0.073, meaning that only 7% of the variation in profit per kilogram can be explained by Total Wage (per day). This value is insignificant, and since the p-value is also greater than 0.05, we do not have sufficient evidence to reject H₀ at the 95% confidence interval. Therefore, it can be concluded that there is no significant relationship between profit per kilogram and Total Wage (per day).
Additionally, it would be unnecessary to derive a Simple Linear Regression (SLR) equation for these variables.
The Economic/Industrial Impact on Daily wage hike to 2000 — Interpretation
Therefore, considering the above correlations, we now have a clearer understanding of how variations in the daily wages of tea plantation laborers impact production cost per kilogram and selling price per kilogram. Additionally, we identified that there is no significant effect on profit per kilogram or total tea production resulting from changes in total wage per day. This implies that these factors are influenced by other variables not considered in this analysis.
Thus, we can now address the question raised earlier: “Can tea companies withstand an increase in the daily wages of tea plantation laborers to LKR 2000, even though they are currently operating at a loss?” The answer is: Yes, they can. The reason for their profits or losses does not appear to be directly linked to wage increases or decreases, as wage variations do not significantly affect profit or loss per kilogram. In simple terms, tea companies need to focus on addressing other significant factors to improve profitability rather than attributing financial challenges to wage adjustments.
Now, let us examine the impact on production cost per kilogram and selling price per kilogram when the total daily wage is increased to LKR 2000.
From the derived equations:
- ŷ (Production Cost per Kg) = 29.8 + 0.64 × Total Wage (Per Day)
- ŷ (Selling Price per Kg) = 24.6 + 0.82 × Total Wage (Per Day)
By substituting LKR 2000 for Total Wage (Per Day), we obtain:
- ŷ (Production Cost per Kg) = 29.8 + 0.64 × 2000 = LKR 1309.8
- ŷ (Selling Price per Kg) = 24.6 + 0.82 × 2000 = LKR 1664.6
This indicates that increasing the daily wages of tea plantation laborers to LKR 2000 would likely result in the highest selling price per kilogram at the Colombo Auction (Gross) observed during the period from 1988 to 2024. However, it may also lead to the highest production cost per kilogram recorded during the same period.
While the increased selling price could potentially allow companies in the tea industry to generate higher profits, this outcome is not guaranteed. It would depend on their ability to optimize other costs and manage operational efficiency. Thus, although there is a possibility for improved profitability with such a wage hike, it remains contingent upon effective cost management and market dynamics.
Social Impact Analysis
For the social impact analysis of increasing the daily wages of tea plantation laborers to LKR 2000, I have developed a NetLogo simulation. The primary objective is to evaluate how this wage increase influences the tea plantation laborers themselves and their social behavior. Conducting such a simulation before implementing the wage increase in society is crucial, as it allows us to observe potential societal responses and outcomes in a controlled environment.
To achieve this, I will incorporate the results obtained from the previous Economic and Industrial Impact Analysis directly into the NetLogo model. Additionally, I will assess socially impactful variables for tea plantation workers (e.g., living wage per month, total labor force in plantations, etc.) in this section, integrating them into the NetLogo simulation if necessary.
NetLogo Simulation Explanation
The NetLogo model world is structured as shown in Figure 9.

Figure 9 — Netlogo Simulation for Social Impact Analysis
This model has three main components: Workers, which are the agents representing plantation laborers; Tea Patches, which represent tea bushes that can be harvested and regrow over time; and the Social Environment, which includes parameters controlling wages, profits, and living costs.
Here, we can set the following parameters,
- Initial Labor Count: (1 unit of labor in the simulation represents 100 laborers in reality)
- Daily_Wage: Base daily wage (target: LKR 2000)
- Profit_kilo: Company profitability per kg (-110 to +100 LKR)
- Working_Days: Monthly working days (typically 20–30)
- Living_Wage: Minimum acceptable monthly income for a tea laborer
- Extra_Pluck_Bonus_kilo: Performance-based bonus system for every additional 1 kg of tea plucked
The worker’s potential income system is calculated as follows.
Total Monthly Income = Base Salary + Performance Bonus
- Base Salary: Daily Wage (input from slider) × Working Days (input from slider)
- Performance Bonus: Extra Bonus (input from slider) × Number of Tea Patches Harvested (calculated by summing the number of “hits” on green patches in the world by the turtle (tea laborer))
- Daily Earnings: Total Monthly Income ÷ Working Days
Example Calculation:
- Daily Wage = 100 LKR (input from slider)
- Working Days = 25 (input from slider)
- Extra Bonus = 10 LKR per kilo (input from slider)
- Worker harvested 15 tea patches
Total Income = (100 × 25) + (10 × 15) = 2,500 + 150 = 2,650 LKR/month Daily Earnings = 2,650 ÷ 25 = 106 LKR/day
Tea Growth Dynamics is as follows.
Green patches = Tea available for picking and When workers walk on green → turns brown (tea harvested). Then Brown patches slowly grow back to green
Growth Rate = 0.1 + (Profit Factor × 0.3)
Where,
- Low profit (-110) = Slow growth (0.1)
- High profit (+100) = Fast growth (0.4)
Meaning: Profitable companies regrow tea faster
- Worker Survival Rules are as follows
Rule A: Must Find Work (Profitability Check)
- Workers wander randomly looking for green tea
- Timer resets every time they find tea
- If no tea found for 30 ticks → FIRED (red color)
Reason: Company doesn’t have enough work
Rule B: Must Earn Enough (Wage Check)
- Every 90 ticks, there’s a “salary review day”
- If Total Income < Living Wage during review → QUIT (magenta color)
Reason: Can’t afford basic living costs
The color code for workers in the simulation is as follows:
- Blue: Normal, adequate pay
- Orange: Warning — income below living wage
- Green: Just successfully harvested tea
- Red: Fired (no work found)
- Magenta: Quit (not paid enough)
The number displayed on each turtle represents the total extra harvest that the worker has plucked, measured in kilograms.
Scenario 1: With 22 Working Days Per Month
Identifying Slider Inputs and Parameter Values for NetLogo Simulation
Here, we will determine the slider inputs and parameter values for the simulation based on the available data and assumptions.
-
Daily_Wage: Since the goal is to increase the daily wages of tea plantation laborers to LKR 2000, the Daily_Wage parameter is set to 2000.
-
Extra_Pluck_Bonus_Kilo: The bonus for extra plucking is given as LKR 50 per kilogram. Therefore, the Extra_Pluck_Bonus_Kilo parameter is set to 50.
-
Initial_Labor_Count: The total tea labor force was 72,135 in 2024, with a wage of LKR 1700. In 2020, when wages were LKR 750, the labor force was 108,157. To estimate the initial labor count for this simulation, we calculate the average labor force over the last five years:
(108,157+99,152+90,146+81,140+72,135)/5=90,146(108,157+99,152+90,146+81,140+72,135)/5 = 90,146
Since 1 unit in the NetLogo simulation represents 100 laborers in real life, the input value is rounded up to 902.
Note: Users can choose to use the current labor force (72,135) instead, depending on their assumptions. For this simulation, the average labor force is used, assuming that the wage increase to LKR 2000 has encouraged previously departed workers to rejoin.
-
Profit_per_Kilo: As determined in the Economic and Industrial Impact Analysis, there is no significant relationship between profit per kilo and other variables. Therefore, the Profit_per_Kilo parameter is set to the 2024 value of -5 (a loss of LKR 5.47 per kilo). Users can adjust this value for different scenarios, but for demonstration purposes, this value is used.
-
Living_Wage:The minimum living wage for 2024 is given as LKR 48,584, with a Consumer Price Index (CPI) of -0.43. To calculate the living wage for 2025, we apply the same CPI: 48,584×(100−0.43)/100=48,37548,584×(100−0.43)/100=48,375 Thus, the Living_Wage parameter is set to 48,375.
-
Working_Days: The total working days in a month are typically considered to be 22. Since these workers receive daily wages, they can work for any number of days (up to 30 or 31). For now, the simulation assumes the usual setting of 22 days per month. If this assumption proves insufficient, the simulation can be rerun with an increased number of working days. For the time being, it is set to 22.
Summary of Parameter Values:
- Daily_Wage: 2000
- Extra_Pluck_Bonus_Kilo: 50
- Initial Labor Count: 902 (NetLogo units, representing 90,146 laborers)
- Profit_per_Kilo: -5
- Living_Wage: 48,375
- Working_Days: 22
Analysis and Interpretation
For Scenario 1, the above parameters were set, and the model was run multiple times, yielding the results shown in Figure 10 below.

Figure 10 — Simulation Result for 22 working days per month
According to the figure, we can observe that for nearly three months (since 1 day = 1 tick), the tea plantations were left unattended by laborers. This was primarily because the laborers could not afford the minimum living wage for three consecutive months. This situation accounts for the majority of cases, around 92%, while the remaining 8% left due to their inability to find work on the plantations. This was likely caused by the continuous losses incurred by plantation companies. To minimize this issue, companies need to generate profits.
However, returning to the main concern, if tea plantation laborers worked only the regular working days (22 days per month), similar to other private or government workers, they might be forced to quit their jobs due to their inability to meet basic needs. If this trend persists, it could negatively impact the tea plantation companies themselves. Additionally, Sri Lanka’s unemployment rate could increase, despite showing a decreasing trend over the years (1988–2024), as illustrated in Figure 11 below.

Figure 11- Unemployment Rate in SL 1988–2024
The reason for their unemployment is that most of the time, these laborers have not acquired other skills that would qualify them for alternative jobs. If they lose their current jobs, they are likely to become unemployed, complicating their lives further.
Scenario 2: With at least 25 Working Days Per Month
Identifying Slider Inputs and Parameter Values for NetLogo Simulation
Since 22 working days per month is not an optimal option, I tested various combinations with 23 and 24 working days while keeping all other parameters constant. However, these scenarios yielded results similar to those observed in Scenario 1. This indicates that tea plantation laborers need to work at least 25 days per month to cover their basic living wages.
To simulate this, I will set all other parameters to remain the same as in Scenario 1, except for the Working_Days, which will be adjusted to 25 days per month. The input parameters for this specific scenario are as follows:
Summary of Parameter Values:
- Daily_Wage: 2000
- Extra_Pluck_Bonus_Kilo: 50
- Initial Labor Count: 902 (NetLogo units, representing 90,146 laborers)
- Profit_per_Kilo: -5
- Living_Wage: 48,375
- Working_Days: 25
The resulting NetLogo world after multiple runs and for a constant time period is as follows (one of the runs is shown in Figure 12 below).

Figure 12 — Simulation Result for 25 working days per month
According to Figure 12, we can observe that no one leaves the plantation due to wage concerns anymore. The only issue now arises when workers are unable to find work, which is caused by plantation companies failing to make a profit and lacking sufficient funds to manage the tea plantations optimally. Additionally, tea plantation workers can now potentially earn even higher wages than LKR 2000 if they exceed the daily “norm” or target for tea plucking. Importantly, since they are working 25 days per month, this means they will often work on Saturdays. However, this is acceptable because they still have 5–6 days of rest per month, which they can use for personal tasks or other activities.
Another interesting pattern captured in the simulation is shown in Figure 13 below.

Figure 13 — Workers departure over time with 25 working days
According to Figure 13, we can see that worker departures become more stabilized over time when the daily wages of tea plantation laborers are increased to LKR 2000, even while holding all other factors constant. This is a positive sign for both the plantation workers themselves and for the tea plantation companies.
The Social Impact on Daily wage hike to 2000 — Interpretation
According to the above scenarios, we can see that increasing the daily wages to LKR 2000 could have a positive social impact on workers. However, they would need to work at least 25 days per month to meet their basic needs. If they work more, they can earn above the living wage, but the downside is that they may not have sufficient time to take care of their personal lives and other responsibilities, as they would essentially be dedicating the entire month to their job.
If wages are increased to LKR 2000, it would enable workers to maintain better social connections and manage their personal lives more effectively. The primary concern, however, lies with the profitability of the tea plantation companies. If these companies can generate profits, they will be better equipped to manage the tea plantations efficiently. This, in turn, would help retain workers in their current jobs and reduce departures caused by issues related to company profitability (e.g., lack of facilities provided by the companies). Such departures currently account for approximately 48%–50% of the initial labor force in the plantations.
Conclusion
Therefore, considering the potential economic, social, and industrial impacts of increasing the daily wages of tea plantation laborers to LKR 2000, we can conclude the following:
There is no significant correlation between wage increases (or decreases) and the profitability of tea plantation companies. Profitability primarily depends on other costs, such as total production costs (all costs associated with tea production except daily wage payments to laborers). If tea companies wish to increase their profits, they need to conduct a root cause analysis and address the underlying issues.
However, if tea plantation companies manage to increase their profitability, it would positively impact the social behavior and stability of tea plantation laborers. This is because laborers are not only concerned about the daily wages they are paid (and obviously, increasing the daily wage to LKR 2000 will encourage them to stay in their jobs), but they also worry about the facilities provided to them.
In conclusion, increasing the daily wage to LKR 2000 would resolve part of the challenges faced by tea plantation workers. However, the remaining issues depend on how effectively tea plantation companies can improve their profitability and provide better working conditions and facilities.
메타데이터
- post_id
- ea7a5b84e55e
- slug
- analyzing-the-impact-of-increasing-daily-wages-of-tea-plantation-labourers-to-lkr-2000-ea7a5b84e55e
- url
- https://medium.com/@chamodmalshankandage/analyzing-the-impact-of-increasing-daily-wages-of-tea-plantation-labourers-to-lkr-2000-ea7a5b84e55e
- canonical_url
- https://medium.com/@chamodmalshankandage/analyzing-the-impact-of-increasing-daily-wages-of-tea-plantation-labourers-to-lkr-2000-ea7a5b84e55e
- author_url
- https://medium.com/@chamodmalshankandage
- status
- ok
- fetched_at
- 2026-07-21 23:55:22