Unweighted, Weighted Aggregate Method

To measure the growth and progress of an economy, economists and scientists use many statistical tools. One such very important tool are index numbers. They help reveal the trends and tendencies of the economy and also help in the formulation of economic policies and laws.

There are broadly three types of index numbers price index numbers, value index numbers, and quantity index numbers.

Very simply put, index numbers help us observe the change in some quantity that we cannot otherwise easily observe or measure. For example, we cannot directly measure the growth of business activity in an economy. However, we can study the changes in factors that influence this business activity.

So an index number is a tool to measure the change in a variable quantity that has happened over a defined period of time. These index numbers are not directly measurable, they are represented as percentages which express the relative changes in quantity.

Quantity Index Numbers

Now we will specifically understand what are quantity index numbers. Quantity index numbers measure the change in the quantity or volume of goods sold, consumed or produced during a given time period. Hence it is a measure of relative changes over a period of time in the quantities of a particular set of goods.

Just like price index numbers and value index numbers, there are also two types of quantity index numbers, namely

  • Unweighted Quantity Indices
  • Weighted Quantity Indices

Let us take a look at the various methods, formulas, and examples of both these types of quantity index numbers.

Unweighted Index: Simple Aggregate Method

Here we do a simple and direct comparison of the aggregate quantities of the current year, with those of the previous year. We express this index number as a percentage. No weights are assigned, it is the simplest calculation. The formula is as follows,

Q01=(ΣQ1/ΣQ0)×100

where, Q1 is the quantity of the current year, and Q0 is the quantity of the previous year,

Unweighted Index: Simple Average of Quantity Method

In this method, we take the aggregate quantities of the current year as a percentage of the quantity of the base year. Then to obtain the index number, we average this percentage figure. So the formula under this method is as follows,

Q01= (ΣQ1/ΣQ0) × 100÷N

where N is the total number of items

Weighted Index: Simple Aggregative Method

There are a few various methods for calculating this index number. We will take a look at some of the most important ones.

1) Laspeyres Method

In this method, the base price is taken as the weight. We only use the price of the base year (P0), not the current year. The formula is as follows,

Q01= (ΣQ1P0/ΣQ0P0) × 100

2) Paasche’s Method

Here, the current year price (P1) of the commodity is taken as the weight.

Q01= (ΣQ1P0/ΣQ0P0) × 100

3) Dorbish & Bowley’s Method

Q01= (ΣQ1P0/ΣQ0P0) + (ΣQ1P1/ΣQ0P1) ÷ 2

Weighted Index: Weighted Average of Relative Method

In this method, we use the arithmetic mean for averaging the values. The formula is a little more complex as seen below,

Q01= ΣQV/ ΣV

where

Q= Σq1/Σq0

and

V=q0p0

Cost of Living Index Number

Uses of cost of living index number:

(i) It is used in wage negotiations, dearness allowance, bonus etc., to the workers.

(ii) The cost of living index number measures the change in the retail prices of a specified quantity of goods and services.

(iii) It is also useful to the government in framing policies relating to wages.

(iv) It is used as measures of change in the purchasing power of money and real income.

The cost-of-living index, or general index, shows the difference in living costs between cities. The cost of living in the base city is always expressed as 100. The cost of living in the destination is then indexed against this number. So to take a simple example, if London is the base (100) and New York is the destination, and the New York index is 120, then New York is 20% more expensive than London. Similarly, if London is the base and Budapest is the destination, and the Budapest index is 80, than the cost of living in Budapest is 80% of London’s.

What’s the methodology behind the index?

The cost-of-living index expresses the difference in the cost of living between any two cities in the survey. How is this index calculated?

Using exactly the same price data, but different methods of calculation, a number of different people could come up with a number of markedly different indices. The challenge, therefore, when seeking to construct an index is to know which method is best for the problem at hand and to represent equitably (in one figure) the general trend of price differences in separate locations. To illustrate this point, let us take a simple price survey comparing two fictional cities, “Mumbai” and “Delhi.”

  Mumbai  Delhi 
Bread (1kg)  1.00  1.25 
Potatoes (1kg)  3.00  2.00 
Coffee (1kg)  2.50  1.75 
Sugar (1kg)  1.00  1.75 
TOTAL  7.50  6.75 

Assuming we give equal weight to each of the products, which of the two towns deserves the higher cost of living index number? The answer is: it all depends on how the calculation is made.

1) Mumbai is more expensive if we simply add up the prices of the four items in the index and compare the two cities on that basis.

2) Delhi, however, is more expensive when we use Mumbai as a base city and calculate an index based on the average of relative prices in the two cities:

  Mumbai  Delhi 
Bread  100  125 
Potatoes  100  67 
Coffee  100  70 
Sugar  100  175 
Index  100  109 

However, if the same calculation is done with Delhi serving as a base city, Mumbai becomes the more expensive city:

  Delhi  Mumbai 
Bread  100  80 
Potatoes  100  150 
Coffee  100  143 
Sugar  100  57 
Index  100  107.50 

Thus with the standard price-relatives calculation we can end up in the paradoxical situation where each city is more expensive than the other.

3) Using a different method, both Delhi and Mumbai would have the same index number, ie 100, and neither would be considered more expensive than the other. Such a calculation would be made according to a well-established statistical formula that takes prices in both cities, makes an average of them, and uses this average as the basis for the index comparison. This formula, adopted by the Economist Intelligence Unit for its indices, has some distinct advantages over the standard price-relatives calculation described in Step 2 above. With the EIU formula, for example, the paradoxical situation of the two cities being more expensive than each other cannot arise: if city A = 100 and city B = 110, then this relationship is maintained, even if city B is used as a base (when B = 100 then A = 91). In other words, the EIU indices are reversible. This property ensures that the cost of living allowances established with the aid of the indices are consistent in that executives transferred from city A to B can be dealt with on the same footing as those transferred from city B to A. In addition, the indices are nearly circular. This means that the relationship between any three cities is maintained regardless of which of the cities is used as a base with which to compare the other two. This logical inter-relationship is important in assuring equitable cost of living compensation as executives are transferred from location to location.

The index formula. The index is based on the arithmetic mean of price levels in the two selected cities. In order to calculate the index for the two hypothetical cities examined on the previous page, we must first calculate the average price of each item:

  Mumbai  Delhi  Average price 
Bread  1.00  1.25  1.125 
Potatoes  3.00  2.00  2.500 
Coffee  2.50  1.75  2.125 
Sugar  1.00  1.75  1.375 

Next we compare prices in each town to these average prices:

  Average  Mumbai  Delhi 
Bread  100  89  111 
Potatoes  100  120  80 
Coffee  100  118  82 
Sugar  100  73  127 
General Index  100  100  100 

As we can see the relationship between Mumbai and Delhi prices remains intact: bread is still 25% more expensive in Delhi, potatoes are still 50% more expensive in Mumbai. If we want to compare Mumbai as a base city to Delhi, we must divide Delhi’s index by that of Mumbai and multiply by 100. The result is 100. If we reverse the operation and use Delhi as base, the result is also 100. The two cities are equally expensive.

Range and co-efficient of Range

The range is a measure of dispersion that represents the difference between the highest and lowest values in a dataset. It provides a simple way to understand the spread of data. While easy to calculate, the range is sensitive to outliers and does not provide information about the distribution of values between the extremes.

Range of a distribution gives a measure of the width (or the spread) of the data values of the corresponding random variable. For example, if there are two random variables X and Y such that X corresponds to the age of human beings and Y corresponds to the age of turtles, we know from our general knowledge that the variable corresponding to the age of turtles should be larger.

Since the average age of humans is 50-60 years, while that of turtles is about 150-200 years; the values taken by the random variable Y are indeed spread out from 0 to at least 250 and above; while those of X will have a smaller range. Thus, qualitatively you’ve already understood what the Range of a distribution means. The mathematical formula for the same is given as:

Range = L – S

where

L: The Largets/maximum value attained by the random variable under consideration

S: The smallest/minimum value.

Properties

  • The Range of a given distribution has the same units as the data points.
  • If a random variable is transformed into a new random variable by a change of scale and a shift of origin as:

Y = aX + b

where

Y: the new random variable

X: the original random variable

a,b: constants.

Then the ranges of X and Y can be related as:

RY = |a|RX

Clearly, the shift in origin doesn’t affect the shape of the distribution, and therefore its spread (or the width) remains unchanged. Only the scaling factor is important.

  • For a grouped class distribution, the Range is defined as the difference between the two extreme class boundaries.
  • A better measure of the spread of a distribution is the Coefficient of Range, given by:

Coefficient of Range (expressed as a percentage) = L – SL + S × 100

Clearly, we need to take the ratio between the Range and the total (combined) extent of the distribution. Besides, since it is a ratio, it is dimensionless, and can, therefore, one can use it to compare the spreads of two or more different distributions as well.

  • The range is an absolute measure of Dispersion of a distribution while the Coefficient of Range is a relative measure of dispersion.

Due to the consideration of only the end-points of a distribution, the Range never gives us any information about the shape of the distribution curve between the extreme points. Thus, we must move on to better measures of dispersion. One such quantity is Mean Deviation which is we are going to discuss now.

Interquartile range (IQR)

The interquartile range is the middle half of the data. To visualize it, think about the median value that splits the dataset in half. Similarly, you can divide the data into quarters. Statisticians refer to these quarters as quartiles and denote them from low to high as Q1, Q2, Q3, and Q4. The lowest quartile (Q1) contains the quarter of the dataset with the smallest values. The upper quartile (Q4) contains the quarter of the dataset with the highest values. The interquartile range is the middle half of the data that is in between the upper and lower quartiles. In other words, the interquartile range includes the 50% of data points that fall in Q2 and

The IQR is the red area in the graph below.

The interquartile range is a robust measure of variability in a similar manner that the median is a robust measure of central tendency. Neither measure is influenced dramatically by outliers because they don’t depend on every value. Additionally, the interquartile range is excellent for skewed distributions, just like the median. As you’ll learn, when you have a normal distribution, the standard deviation tells you the percentage of observations that fall specific distances from the mean. However, this doesn’t work for skewed distributions, and the IQR is a great alternative.

I’ve divided the dataset below into quartiles. The interquartile range (IQR) extends from the low end of Q2 to the upper limit of Q3. For this dataset, the range is 21 – 39.

Quartiles, Quartile Deviation and Quartile co-efficient

The Quartile Deviation is a simple way to estimate the spread of a distribution about a measure of its central tendency (usually the mean). So, it gives you an idea about the range within which the central 50% of your sample data lies. Consequently, based on the quartile deviation, the Coefficient of Quartile Deviation can be defined, which makes it easy to compare the spread of two or more different distributions. Since both of these topics are based on the concept of quartiles, we’ll first understand how to calculate the quartiles of a dataset before working with the direct formulae.

Quartiles

A median divides a given dataset (which is already sorted) into two equal halves similarly, the quartiles are used to divide a given dataset into four equal halves. Therefore, logically there should be three quartiles for a given distribution, but if you think about it, the second quartile is equal to the median itself! We’ll deal with the other two quartiles in this section.

  • The first quartileor the lower quartile or the 25th percentile, also denoted by Q1corresponds to the value that lies halfway between the median and the lowest value in the distribution (when it is already sorted in the ascending order). Hence, it marks the region which encloses 25% of the initial data.
  • Similarly, the third quartileor the upper quartile or 75th percentile, also denoted by Q3, corresponds to the value that lies halfway between the median and the highest value in the distribution (when it is already sorted in the ascending order). It, therefore, marks the region which encloses the 75% of the initial data or 25% of the end data.

For a better understanding, look at the representation below for a Gaussian Distribution:

The Quartile Deviation

Formally, the Quartile Deviation is equal to the half of the Inter-Quartile Range and thus we can write it as:

Qd=(Q3–Q1)/2

Therefore, we also call it the Semi Inter-Quartile Range.

  • The Quartile Deviation doesn’t take into account the extreme points of the distribution. Thus, the dispersion or the spread of only the central 50% data is considered.
  • If the scale of the data is changed, the Qd also changes in the same ratio.
  • It is the best measure of dispersion for open-ended systems (which have open-ended extreme ranges).
  • Also, it is less affected by sampling fluctuations in the dataset as compared to the range (another measure of dispersion).
  • Since it is solely dependent on the central values in the distribution, if in any experiment, these values are abnormal or inaccurate, the result would be affected drastically.

The Coefficient of Quartile Deviation

Based on the quartiles, a relative measure of dispersion, known as the Coefficient of Quartile Deviation, can be defined for any distribution. It is formally defined as:

Coefficient of Quartile Deviation = {(Q3–Q1)/(Q3+Q1)}×100

Since it involves a ratio of two quantities of the same dimensions, it is unit-less. Thus, it can act as a suitable parameter for comparing two or more different datasets which may or may not involve quantities with the same dimensions.

So, now let’s go through the solved examples below to get a better idea of how to apply these concepts to various distributions.

Mean deviation with mean, Co-efficient of mean deviation

To understand the dispersion of data from a measure of central tendency, we can use mean deviation. It comes as an improvement over the range. It basically measures the deviations from a value. This value is generally mean or median. Hence although mean deviation about mode can be calculated, mean deviation about mean and median are frequently used.

Note that the deviation of an observation from a value a is d= x-aTo find out mean deviation we need to take the mean of these deviations. However, when this value of a is taken as mean, the deviations are both negative and positive since it is the central value.

This further means that when we sum up these deviations to find out their average, the sum essentially vanishes. Thus to resolve this problem we use absolute values or the magnitude of deviation. The basic formula for finding out mean deviation is :

Mean deviation= Sum of absolute values of deviations from ‘a’ ÷ The number of observations

Coefficient of Mean Deviation:

It is calculated to compare the data of two series. The coefficient of mean deviation is calculated by dividing mean deviation by the average. If deviations are taken from mean, we divide it by mean, if the deviations are taken from median, then it is divided by mode and if the “deviations are taken from median, then we divide mean deviation by median.

  1. For Discrete Series:

M.D. = ∑fdy/N; Where; N=∑f

And dy is the deviation of variable from X, M or Z ignoring signs (Taking +ive signs only).

Steps to Calculate:

  1. Take X, M or Z series as desired.
  2. Take deviations ignoring signs.
  3. Multiply dy by respective f; get ∑fdy
  4. Use the following formula

M.D. = ∑fdy/N

(Note : If value of X or M or Z is in decimal fractions better use Direct Method to get result easily)

When Mean or Median or Mode is in Fractions, in that case, Direct formula is applied

  1. For Continuous Series:

For Continuous Series also ;

M.D. = fdy/N

Standard deviation with co-efficient of Variance

As the name suggests, this quantity is a standard measure of the deviation of the entire data in any distribution. Usually represented by or σ. It uses the arithmetic mean of the distribution as the reference point and normalizes the deviation of all the data values from this mean.

Therefore, we define the formula for the standard deviation of the distribution of a variable X with n data points as:

Variance

Another statistical term that is related to the distribution is the variance, which is the standard deviation squared (variance = SD² ). The SD may be either positive or negative in value because it is calculated as a square root, which can be either positive or negative. By squaring the SD, the problem of signs is eliminated. One common application of the variance is its use in the F-test to compare the variance of two methods and determine whether there is a statistically significant difference in the imprecision between the methods.

In many applications, however, the SD is often preferred because it is expressed in the same concentration units as the data. Using the SD, it is possible to predict the range of control values that should be observed if the method remains stable. As discussed in an earlier lesson, laboratorians often use the SD to impose “gates” on the expected normal distribution of control values.

Coefficient of Variation

Another way to describe the variation of a test is calculate the coefficient of variation, or CV. The CV expresses the variation as a percentage of the mean, and is calculated as follows:

CV% = (SD/Xbar)100

In the laboratory, the CV is preferred when the SD increases in proportion to concentration. For example, the data from a replication experiment may show an SD of 4 units at a concentration of 100 units and an SD of 8 units at a concentration of 200 units. The CVs are 4.0% at both levels and the CV is more useful than the SD for describing method performance at concentrations in between. However, not all tests will demonstrate imprecision that is constant in terms of CV. For some tests, the SD may be constant over the analytical range.

The CV also provides a general “feeling” about the performance of a method. CVs of 5% or less generally give us a feeling of good method performance, whereas CVs of 10% and higher sound bad. However, you should look carefully at the mean value before judging a CV. At very low concentrations, the CV may be high and at high concentrations the CV may be low. For example, a bilirubin test with an SD of 0.1 mg/dL at a mean value of 0.5 mg/dL has a CV of 20%, whereas an SD of 1.0 mg/dL at a concentration of 20 mg/dL corresponds to a CV of 5.0%.

Skewness

Skewness is a statistical measure that indicates the degree and direction of asymmetry in a frequency distribution. When data is distributed evenly around the central value, the distribution is said to be symmetrical. However, if one side of the distribution extends farther than the other, the distribution is skewed.

In Business Statistics, skewness helps researchers and managers understand the nature of data distribution, identify trends, and make informed decisions. It is commonly used in the analysis of income, profits, wages, sales, investment returns, and market behavior.

Definition of Skewness

Skewness refers to the extent to which a distribution deviates from symmetry. It measures whether the observations are concentrated more on one side of the distribution than the other.

A distribution may be:

  • Symmetrical
  • Positively Skewed
  • Negatively Skewed

Types of Skewness

1. Symmetrical Distribution

A symmetrical distribution has equal frequencies on both sides of the central value.

Characteristics

  • Mean = Median = Mode
  • No skewness
  • Skewness coefficient = 0

Example: The distribution of heights of a large group of people often approximates a symmetrical distribution.

Diagram

2. Positive Skewness (Right Skewness)

A distribution is positively skewed when the tail extends toward the right side.

Characteristics

  • Mean > Median > Mode
  • More observations are concentrated at lower values.
  • A few high values pull the mean to the right.

Example: Income distribution in many countries where a small number of people earn very high incomes.

Diagram

3. Negative Skewness (Left Skewness)

A distribution is negatively skewed when the tail extends toward the left side.

Characteristics

  • Mean < Median < Mode
  • More observations are concentrated at higher values.
  • A few low values pull the mean to the left.

Example: Marks obtained in an easy examination where most students score high marks.

Diagram

Importance of Skewness

  • Helps Understand the Nature of Data Distribution

Skewness helps statisticians and business analysts understand whether a dataset is symmetrical or asymmetrical. It reveals the direction and degree of deviation from a normal distribution. By examining skewness, researchers can identify whether observations are concentrated toward higher or lower values. This understanding is essential for interpreting data accurately. In business statistics, knowing the nature of distribution helps managers evaluate performance, customer behavior, and market trends more effectively, leading to better analysis and decision-making.

  • Assists in Business Decision-Making

Business decisions often depend on accurate interpretation of statistical data. Skewness provides valuable insights into the distribution of sales, profits, costs, and customer preferences. By understanding whether data is positively or negatively skewed, managers can identify unusual patterns and take appropriate actions. It helps in resource allocation, strategic planning, and performance evaluation. Therefore, skewness serves as an important analytical tool that supports informed and rational decision-making in various business activities and organizational operations.

  • Useful in Forecasting and Planning

Forecasting future trends requires a proper understanding of past and present data. Skewness helps identify the distribution pattern of historical observations, enabling analysts to make more accurate predictions. If data is highly skewed, forecasting models may need adjustments to improve reliability. Businesses use skewness while planning production, inventory, marketing strategies, and financial investments. By understanding the direction of data concentration, organizations can anticipate future developments and prepare suitable plans, reducing uncertainty and improving operational efficiency.

  • Helps in Selecting Appropriate Statistical Methods

Many statistical techniques assume that data follows a normal or symmetrical distribution. Skewness helps determine whether these assumptions are valid. If a dataset is highly skewed, analysts may need to use alternative methods or transform the data before analysis. This ensures the accuracy and validity of statistical results. In research and business studies, selecting the correct analytical technique is crucial for drawing reliable conclusions. Therefore, skewness plays an important role in choosing suitable statistical tools and procedures.

  • Identifies the Presence of Extreme Values

Skewness helps detect the influence of extreme values or outliers in a dataset. A highly skewed distribution often indicates that a few observations are significantly larger or smaller than the majority. Identifying such values is important because they can affect averages, forecasts, and business decisions. Managers and researchers can investigate these unusual observations to determine whether they represent genuine trends or data errors. Thus, skewness contributes to more accurate data interpretation and enhances the quality of statistical analysis.

  • Useful in Financial and Investment Analysis

In finance, skewness is widely used to analyze investment returns, stock prices, and financial risks. Investors prefer to understand whether returns are concentrated around gains or losses. Positive and negative skewness provide information about potential opportunities and risks associated with investments. Financial analysts use skewness to evaluate portfolio performance and make informed investment decisions. Therefore, skewness is an important measure in risk assessment, helping businesses and investors manage uncertainty and improve financial planning.

  • Facilitates Comparison of Different Distributions

Skewness enables comparison between different datasets by showing the direction and degree of asymmetry. Two datasets may have similar averages but differ significantly in their distribution patterns. By measuring skewness, analysts can identify these differences and gain deeper insights into the data. Businesses often compare sales performance, customer behavior, employee productivity, and financial results using skewness measures. This comparative analysis helps managers understand relative performance and make more effective decisions based on statistical evidence.

  • Enhances Research and Market Analysis

Skewness is an important tool in research and market analysis because it provides information about consumer behavior, market demand, and economic conditions. Researchers use skewness to study patterns and identify trends within datasets. In marketing, understanding skewed distributions helps businesses segment customers and develop targeted strategies. It also assists in evaluating survey results and market responses. By offering a clearer picture of data behavior, skewness improves the quality of research findings and supports better business and policy decisions.

Limitations of Skewness

  • Highly Sensitive to Extreme Values

One of the major limitations of skewness is its sensitivity to extreme values or outliers. A few unusually large or small observations can significantly influence the skewness coefficient and create a misleading impression of the distribution. In business data, unusual sales figures, profits, or losses may distort the measure of skewness. As a result, the calculated value may not accurately represent the majority of observations. Therefore, analysts must carefully examine the presence of outliers before interpreting skewness and drawing conclusions from statistical data.

  • Does Not Measure Dispersion

Skewness measures only the asymmetry of a distribution and provides no information about the spread or variability of data. Two datasets may have the same skewness value but differ greatly in their dispersion. To understand the complete nature of a distribution, skewness must be used along with measures such as range, variance, and standard deviation. Relying solely on skewness can lead to incomplete analysis. Therefore, it should be considered as one aspect of statistical description rather than a comprehensive measure of data characteristics.

  • Different Methods May Give Different Results

There are several methods of measuring skewness, including Karl Pearson’s, Bowley’s, and Kelly’s coefficients. These methods are based on different statistical concepts and may produce different values for the same dataset. Such variations can create confusion in interpretation and comparison. Analysts may find it difficult to determine which measure best represents the distribution. Consequently, the existence of multiple methods reduces the uniformity of skewness measurement and sometimes complicates statistical analysis, especially when comparing results from different studies or datasets.

  • Difficult to Interpret Precisely

Although skewness indicates the direction and degree of asymmetry, its exact interpretation is often difficult. A positive or negative value shows the direction of skewness, but understanding the practical significance of a particular value may not be straightforward. For example, determining whether a skewness coefficient indicates moderate or severe asymmetry requires additional judgment. This complexity may create challenges for managers, researchers, and students. Therefore, skewness values should be interpreted carefully and in conjunction with graphical analysis and other statistical measures.

  • Not Reliable for Small Samples

Skewness may not provide reliable results when calculated from small samples. In small datasets, a few observations can greatly influence the measure, making it unstable and less representative of the population. Sampling fluctuations may cause skewness values to vary considerably from one sample to another. As a result, conclusions based on skewness from limited data may be misleading. For accurate interpretation, larger datasets are generally preferred. Therefore, analysts should exercise caution when using skewness to evaluate distributions based on small samples.

  • Cannot Fully Describe Distribution Shape

Skewness provides information only about asymmetry and does not fully describe the shape of a distribution. Other characteristics, such as kurtosis, modality, and dispersion, are also important for understanding data behavior. Two distributions may have identical skewness values but differ significantly in other aspects. Consequently, skewness alone cannot provide a complete picture of the dataset. Analysts must combine it with additional statistical measures and graphical tools to gain a thorough understanding of the distribution and make informed decisions.

  • Requires Accurate Data

The accuracy of skewness depends heavily on the quality of the data used. Errors in data collection, recording, classification, or tabulation can affect the calculated skewness coefficient and lead to incorrect conclusions. In business statistics, inaccurate sales, profit, or customer data may distort the measure of asymmetry. Therefore, reliable and properly verified data is essential for meaningful skewness analysis. This dependence on data accuracy represents a limitation because errors at any stage of data handling can reduce the usefulness of skewness measurements.

  • Limited Use When Used Alone

Skewness has limited usefulness when considered in isolation. While it provides information about asymmetry, it does not explain other important characteristics of the dataset. Effective statistical analysis requires the use of multiple measures, including averages, dispersion, and correlation. If skewness is used alone, analysts may overlook critical aspects of data behavior. Therefore, it should be regarded as a supplementary measure rather than a complete analytical tool. Combining skewness with other statistical techniques leads to more accurate interpretations and better decision-making.

Kurtosis

Kurtosis is a statistical measure that describes the degree of peakedness or flatness of a frequency distribution in comparison with a normal distribution. It indicates how observations are concentrated around the mean and how the tails of the distribution behave.

In Business Statistics, kurtosis helps analysts understand the shape of a distribution and identify whether data contains extreme observations. It is widely used in finance, economics, market research, quality control, and risk analysis.

Definition of Kurtosis

Kurtosis is the measure of the shape of a distribution that indicates the extent to which observations cluster around the center and the thickness of the tails relative to a normal distribution.

The term Kurtosis was introduced by Karl Pearson.

Excess Kurtosis

An excess kurtosis is a metric that compares the kurtosis of a distribution against the kurtosis of a normal distribution. The kurtosis of a normal distribution equals 3. Therefore, the excess kurtosis is found using the formula below:

Excess Kurtosis = Kurtosis – 3

Types of Kurtosis

The types of kurtosis are determined by the excess kurtosis of a particular distribution. The excess kurtosis can take positive or negative values as well, as values close to zero.

1. Mesokurtic

Mesokurtic Distribution is a distribution that has the same degree of peakedness and tail thickness as a normal distribution. It serves as the standard or benchmark against which other types of kurtosis are compared. In a mesokurtic distribution, observations are moderately concentrated around the mean, and the tails are neither too heavy nor too light. The coefficient of kurtosis (β₂) is equal to 3, while excess kurtosis is 0. Many natural and social phenomena approximately follow a mesokurtic pattern. This type of distribution indicates a balanced spread of data without an unusual concentration of extreme values. In business statistics, mesokurtic distributions are often considered ideal because they reflect a normal and predictable pattern of observations.

Example: The distribution of examination scores in a large class often approximates a mesokurtic distribution.

2. Leptokurtic

Leptokurtic Distribution is more peaked than a normal distribution and has heavier tails. In this type of distribution, a large number of observations are concentrated near the mean, while the tails contain more extreme values than a normal distribution. The coefficient of kurtosis (β₂) is greater than 3, and excess kurtosis is positive. Because of its heavy tails, a leptokurtic distribution indicates a higher probability of extreme observations occurring. This characteristic is particularly important in finance and investment analysis, where sudden gains or losses may occur. In business statistics, leptokurtic distributions are useful for identifying situations involving high risk and volatility. The presence of a sharp peak and heavy tails suggests that observations cluster around the center but occasionally produce significant deviations from the average.

Example: Stock market returns often follow a leptokurtic distribution because extreme gains and losses occur more frequently than expected under a normal distribution.

3. Platykurtic

Platykurtic Distribution is flatter than a normal distribution and has lighter tails. In this type of distribution, observations are more evenly spread across the range of data, resulting in a broad and low central peak. The coefficient of kurtosis (β₂) is less than 3, while excess kurtosis is negative. Because the tails are lighter, extreme observations occur less frequently than in a normal distribution. A platykurtic distribution indicates greater dispersion and lower concentration of observations around the mean. In business statistics, such distributions may occur when data is uniformly distributed across different categories. The flatter shape suggests that observations are widely dispersed and that the likelihood of unusually high or low values is relatively small.

Example: The distribution of customer arrivals spread evenly throughout a day may exhibit a platykurtic pattern.

Karl Pearson and Spearman Rank Correlation

Karl Pearson Coefficient of Correlation

Karl Pearson Coefficient of Correlation (also called the Pearson correlation coefficient or Pearson’s r) is a measure of the strength and direction of the linear relationship between two variables. It ranges from -1 to +1, where +1 indicates a perfect positive linear relationship, -1 indicates a perfect negative linear relationship, and 0 indicates no linear relationship. The formula for Pearson’s r is calculated by dividing the covariance of the two variables by the product of their standard deviations. It is widely used in statistics to analyze the degree of correlation between paired data.

The following are the main properties of correlation.

1. Coefficient of Correlation lies between -1 and +1:

The coefficient of correlation cannot take value less than -1 or more than one +1. Symbolically,

-1<=r<= + 1 or | r | <1.

2. Coefficients of Correlation are independent of Change of Origin:

This property reveals that if we subtract any constant from all the values of X and Y, it will not affect the coefficient of correlation.

3. Coefficients of Correlation possess the property of symmetry:

The degree of relationship between two variables is symmetric as shown below:

4. Coefficient of Correlation is independent of Change of Scale:

This property reveals that if we divide or multiply all the values of X and Y, it will not affect the coefficient of correlation.

5. Co-efficient of correlation measures only linear correlation between X and Y.

6. If two variables X and Y are independent, coefficient of correlation between them will be zero.

Karl Pearson’s Coefficient of Correlation is widely used mathematical method wherein the numerical expression is used to calculate the degree and direction of the relationship between linear related variables.

Pearson’s method, popularly known as a Pearsonian Coefficient of Correlation, is the most extensively used quantitative methods in practice. The coefficient of correlation is denoted by “r”.

If the relationship between two variables X and Y is to be ascertained, then the following formula is used:

Properties of Coefficient of Correlation

  • The value of the coefficient of correlation (r) always lies between±1. Such as:r = +1, perfect positive correlation

    r = -1, perfect negative correlation

    r = 0, no correlation

  • The coefficient of correlation is independent of the origin and scale.By origin, it means subtracting any non-zero constant from the given value of X and Y the vale of “r” remains unchanged. By scale it means, there is no effect on the value of “r” if the value of X and Y is divided or multiplied by any constant.
  • The coefficient of correlation is a geometric mean of two regression coefficient. Symbolically it is represented as:
  • The coefficient of correlation is “ zero” when the variables X and Y are independent. But, however, the converse is not true.

Assumptions of Karl Pearson’s Coefficient of Correlation

  • The relationship between the variables is “Linear”, which means when the two variables are plotted, a straight line is formed by the points plotted.
  • There are a large number of independent causes that affect the variables under study so as to form a Normal Distribution. Such as, variables like price, demand, supply, etc. are affected by such factors that the normal distribution is formed.
  • The variables are independent of each other.                                     

Note: The coefficient of correlation measures not only the magnitude of correlation but also tells the direction. Such as, r = -0.67, which shows correlation is negative because the sign is “-“ and the magnitude is 0.67.

Spearman Rank Correlation

Spearman rank correlation is a non-parametric test that is used to measure the degree of association between two variables.  The Spearman rank correlation test does not carry any assumptions about the distribution of the data and is the appropriate correlation analysis when the variables are measured on a scale that is at least ordinal.

The Spearman correlation between two variables is equal to the Pearson correlation between the rank values of those two variables; while Pearson’s correlation assesses linear relationships, Spearman’s correlation assesses monotonic relationships (whether linear or not). If there are no repeated data values, a perfect Spearman correlation of +1 or −1 occurs when each of the variables is a perfect monotone function of the other.

Intuitively, the Spearman correlation between two variables will be high when observations have a similar (or identical for a correlation of 1) rank (i.e. relative position label of the observations within the variable: 1st, 2nd, 3rd, etc.) between the two variables, and low when observations have a dissimilar (or fully opposed for a correlation of −1) rank between the two variables.

The following formula is used to calculate the Spearman rank correlation:

ρ = Spearman rank correlation

di = the difference between the ranks of corresponding variables

n = number of observations

Assumptions

The assumptions of the Spearman correlation are that data must be at least ordinal and the scores on one variable must be monotonically related to the other variable.

Least Square Method

The least square method is the process of finding the best-fitting curve or line of best fit for a set of data points by reducing the sum of the squares of the offsets (residual part) of the points from the curve. During the process of finding the relation between two variables, the trend of outcomes are estimated quantitatively. This process is termed as regression analysis. The method of curve fitting is an approach to regression analysis. This method of fitting equations which approximates the curves to given raw data is the least square.

It is quite obvious that the fitting of curves for a particular data set are not always unique. Thus, it is required to find a curve having a minimal deviation from all the measured data points. This is known as the best-fitting curve and is found by using the least-squares method.

Least Square Method

The least-squares method is a crucial statistical method that is practised to find a regression line or a best-fit line for the given pattern. This method is described by an equation with specific parameters. The method of least squares is generously used in evaluation and regression. In regression analysis, this method is said to be a standard approach for the approximation of sets of equations having more equations than the number of unknowns.

The method of least squares actually defines the solution for the minimization of the sum of squares of deviations or the errors in the result of each equation. Find the formula for sum of squares of errors, which help to find the variation in observed data.

The least-squares method is often applied in data fitting. The best fit result is assumed to reduce the sum of squared errors or residuals which are stated to be the differences between the observed or experimental value and corresponding fitted value given in the model.

There are two basic categories of least-squares problems:

  • Ordinary or linear least squares
  • Nonlinear least squares

These depend upon linearity or nonlinearity of the residuals. The linear problems are often seen in regression analysis in statistics. On the other hand, the non-linear problems generally used in the iterative method of refinement in which the model is approximated to the linear one with each iteration.

Least Square Method Graph

In linear regression, the line of best fit is a straight line as shown in the following diagram:

The given data points are to be minimized by the method of reducing residuals or offsets of each point from the line. The vertical offsets are generally used in surface, polynomial and hyperplane problems, while perpendicular offsets are utilized in common practice.

Least Square Method Formula

The least-square method states that the curve that best fits a given set of observations, is said to be a curve having a minimum sum of the squared residuals (or deviations or errors) from the given data points. Let us assume that the given points of data are (x1,y1), (x2,y2), (x3,y3), …, (xn,yn) in which all x’s are independent variables, while all y’s are dependent ones. Also, suppose that f(x) be the fitting curve and d represents error or deviation from each given point.

Now, we can write:

d1 = y1 − f(x1)

d2 = y2 − f(x2)

d3 = y3 − f(x3)

…..

dn = yn – f(xn)

The least-squares explain that the curve that best fits is represented by the property that the sum of squares of all the deviations from given values must be minimum. i.e:

Sum = Minimum Quantity

Limitations for Least-Square Method

The least-squares method is a very beneficial method of curve fitting. Despite many benefits, it has a few shortcomings too. One of the main limitations is discussed here.

In the process of regression analysis, which utilizes the least-square method for curve fitting, it is inevitably assumed that the errors in the independent variable are negligible or zero. In such cases, when independent variable errors are non-negligible, the models are subjected to measurement errors. Therefore, here, the least square method may even lead to hypothesis testing, where parameter estimates and confidence intervals are taken into consideration due to the presence of errors occurring in the independent variables.

error: Content is protected !!