Symmetrical and Skewed Distributions

A symmetric distribution is one where the left and right hand sides of the distribution are roughly equally balanced around the mean. The histogram below shows a typical symmetric distribution.

For symmetric distributions, the mean is approximately equal to the median. The tails of the distribution are the parts to the left and to the right, away from the mean. The tail is the part where the counts in the histogram become smaller. For a symmetric distribution, the left and right tails are equally balanced, meaning that they have about the same length.

Symmetrical distribution occurs when the values of variables occur at regular frequencies and the mean, median and mode occur at the same point. In graph form, symmetrical distribution often appears as a bell curve. If a line were drawn dissecting the middle of the graph, it would show two sides that mirror each other. Symmetrical distribution is a core concept in technical trading as the price action of an asset is assumed to fit a symmetrical distribution curve over time.

Symmetrical distribution is used by traders to establish the value area for a stock, currency or commodity on a set time frame. This time frame is can be intraday, such as 30 minute intervals, or it can be longer-term using sessions or even weeks and months. A bell curve can be drawn around the price points hit during that time period and it is expected that most of the price action – approximately 68% of price points – will fall within one standard deviation of the centre of the curve. The curve is applied to the y-axis (price) as it is the variable whereas time throughout the period is simply linear. So the area within one standard deviation of the mean is the value area where price and the actual value of the asset are most closely matched.

If the price action takes the asset price out of the value area, then it suggests that price and value are out of alignment. If the breach is to the bottom of the curve, the asset is considered to be undervalued. If it is to the top of the curve, the asset is to be overvalued. The assumption is that the asset will revert to the mean over time.

An Example of How Symmetrical Distribution is Used

Symmetrical distribution is most often used to put price action into context. The further the price action wanders from the value area one standard deviation on each side of the mean, the greater the probability that the underlying asset is being under or overvalued by the market. This observation will suggest potential trades to place based on how far the price action has wandered from the mean for the time period being used. On larger time scales, however, there is a much greater risk of missing the actual entry and exit points.

  • Symmetrical distribution can refer to a bell curve or any curve where a halving line produces mirror images.
  • When traders speak of reversion to the mean, they are referring to the symmetrical distribution of price action overtime.
  • The opposite of symmetrical distribution is asymmetrical distribution, which is a curve that exhibits skewness.

Skewed

A distribution that is skewed right (also known as positively skewed) is shown below.

Now the picture is not symmetric around the mean anymore. For a right skewed distribution, the mean is typically greater than the median. Also notice that the tail of the distribution on the right hand (positive) side is longer than on the left hand side.

From the box and whisker diagram we can also see that the median is closer to the first quartile than the third quartile. The fact that the right hand side tail of the distribution is longer than the left can also be seen.

A distribution that is skewed left has exactly the opposite characteristics of one that is skewed right:

  • The mean is typically less than the median;
  • The tail of the distribution is longer on the left hand side than on the right hand side; and
  • The median is closer to the third quartile than to the first quartile.

Measures of Dispersion Meaning, Absolute and Relative

Measures of dispersion refer to statistical tools used to describe the spread or variability of a dataset. These measures help in understanding the extent to which data points differ from the central tendency (mean, median, or mode). Common measures of dispersion include:

  • Range: The difference between the highest and lowest values.
  • Variance: The average squared deviation of each data point from the mean.
  • Standard deviation: The square root of variance, providing a more interpretable measure of spread.
  • Interquartile range (IQR): The range between the 25th and 75th percentiles.

Characteristics of Measures of Dispersion:

  • A measure of dispersion should be rigidly defined
  • It must be easy to calculate and understand
  • Not affected much by the fluctuations of observations
  • Based on all observations

Classification of Measures of Dispersion

The measure of dispersion is categorized as:

(i) An absolute measure of dispersion:

  • The measures which express the scattering of observation in terms of distances i.e., range, quartile deviation.
  • The measure which expresses the variations in terms of the average of deviations of observations like mean deviation and standard deviation.

(ii) A relative measure of dispersion:

We use a relative measure of dispersion for comparing distributions of two or more data set and for unit free comparison. They are the coefficient of range, the coefficient of mean deviation, the coefficient of quartile deviation, the coefficient of variation, and the coefficient of standard deviation.

Coefficient of Dispersion

Whenever we want to compare the variability of the two series which differ widely in their averages. Also, when the unit of measurement is different. We need to calculate the coefficients of dispersion along with the measure of dispersion. The coefficients of dispersion (C.D.) based on different measures of dispersion are

  • Based on Range = (X max – X min) ⁄ (X max + X min).
  • C.D. based on quartile deviation = (Q3 – Q1) ⁄ (Q3 + Q1).
  • Based on mean deviation = Mean deviation/average from which it is calculated.
  • For Standard deviation = S.D. ⁄ Mean

Coefficient of Variation

100 times the coefficient of dispersion based on standard deviation is the coefficient of variation (C.V.).

C.V. = 100 × (S.D. / Mean) = (σ/ȳ ) × 100

Probable error

Probable Error is basically the correlation coefficient that is fully responsible for the value of the coefficients and its accuracy.

As mentioned, probable error is the coefficient of correlation that supports in finding out about the accurate values of the coefficients. It also helps in determining the reliability of the coefficient.

The calculation of the correlation coefficient usually takes place from the samples. These samples are in pairs. The pairs generally come from a very large population. It is quite an easy job to find out about the limits and bounds of the correlation coefficient.

The correlation coefficient for a population is usually based on the knowledge and the sample relating to the correlation coefficient. Therefore, probable error is the easy way to find out or obtain the correlation coefficient of any population. Hence, the definition is:

Probable Error = 0.674 ×

Here, r = correlation coefficient of ‘n’ pairs of observations for any random sample and N = Total number of observations.

About the Values

  • There is hardly any correlation between the different variables if the value of ‘r’ turns out to be less than the value of the probable error
  • The value of correlation coefficient is generally certain if and only if the value of ‘r’ is around 6 times more than the value of the error.
  •  The value of the probable error is in the bounds -1 and +1(-1≤r≤1). So, we can express it in the following manner.

Probable Limit

To get the upper limit and the lower limit, all we need to do is respectively add and subtract the value of probable error from the value of ‘r.’ This is exactly where the value of correlation of coefficient lies.

ρ (rho) = r ± P.E.

Here, the value of rho is nothing but the correlation coefficient of a population. This is also the limit of the correlation of coefficient. Alongside,

Probable Error = 2/3 SE

Here, S.E is Standard Error of Correlation Coefficient

Standard Error = (1-r2)/√N

Standard Error is basically the standard deviation of any mean. It is the sampling distribution of the standard deviation. The standard error is generally used to refer to any sort of estimate belonging to the standard deviation. Therefore, we use probable error to calculate and check the reliability associated with the coefficient.

Advantages of Standard Error

  • It helps in finding and reducing the sample errors as well as the measurement errors.
  • The standard error of any mean tells about the accuracy of the estimate clearly enough.

Formulas for Calculating Probable Error

Generally, there are three formulas using which we can calculate the probable error. The very first formula is the most common formula to calculate P.E. We use the Pearson product-moment method for calculating the same. It is:

P.E r  product-moment = 0.6745(1-r2)/√N

The second formula is applicable when we need the probable error for rho. We use the Spearman method to calculate the value. The formula for the same is:

P.E. ρ = 0.6745(1-ρ2)/√N {1 + 1.086ρ+ 0.13ρ+ .002ρ6}

The third formula is applicable to the Pearson coefficient ‘r.’  We calculate it through ρ by using the transmutation formula. The value is r = 2 sin (πρ/6). The formula is given by:

  1. E  rfound from ρ = 0.7063 (1 – r2)√N {1 + 1.042r+ 0.008r+ .002r6}

Note: The formula that we are using to calculate probable error is valid and applicable if the given population is normal.

Conditions to find Probable Error

We can find the probable error if and only if the given below conditions are taken care of.

  • The data that we have must be a bell-shaped curve. This means that the data has to give us a normal frequency curve
  • It is important to take the probable error for measuring the statistics from the sample only
  • It is compulsory that the sample items are taken off in an unbiased manner and must remain independent of each other’s value

Simple Aggregative Method

We use this method of construction for computation of index price. As a result, the total cost of any commodity in any given year to the total cost of any commodity in the base year is in percentage form.

Simple Aggregative Price Index – (∑ Pn/ ∑ P0) * 100

Where

∑Pn = Sum of the price of all the respective commodity in the current time period.
∑P= Sum of the price of all the respective commodity in the base period.

The simple aggregative index is very simple to understand. However, there is a serious defect in this method. The first commodity, here, has more influence than the rest two. This is so because the first commodity has a high price than the rest.

Furthermore, if we anyhow change the units, the index number will also go through a change. This is one of the biggest flaws of this methods. Use of absolute quantities turn the tables around. Therefore, considering independent values for the three years would be a better option.

To construct a simple price index, compute the price relatives and average them. Add the price relatives and divide them by the number of items. Table illustrates the construction of a simple index of wholesale prices.

Commodity Prices in 1970(P0) Base

1970=100

Prices in 1980(P1) = P1/P0xl00 Price Relatives

(R)

A Rs . 20 per kg 100 Rs. 25 125
В 5 per kg 100 10 200
С 15 per metre 100 30 200
D 25 per kg 100 30 120
E 200 per quantal 100 450 225
N = 5 500 ∑R = 870

Price index in 1980 = Prices in 1980 / Prices in 1970 x 100

Or ∑P1/P0 x 100 = 870/500 x 100 = 174

Using arithmetic mean, price index in 1980 = ∑R/N = 870/5 = 174

The preceding table shows that 1970 is the base period and 1980 is the year for which the price index has been constructed on the basis of price relatives. The index of wholesale prices in 1980 comes to 174. This means that the price level rose by 74 per cent in 1980 over 1970.

Constructing Index Numbers

An index number is a statistical tool used to measure changes in the value of money. It indicates the average price level of a selected group of commodities at a specific point in time compared to the average price level of the same group at another time.

It represents the average of various items expressed in different units. Additionally, an index number reflects the overall increase or decrease in the average prices of the group being studied. For example, if the Consumer Price Index rises from 100 in 1980 to 150 in 1982, it indicates a 50 percent rise in the prices of the commodities included. Furthermore, an index number shows the degree of change in the value of money (or the price level) over time, based on a chosen base year. If the base year is 1970, we can evaluate the change in the average price level for both earlier and later years.

Construction of Index Number:

1. Define the Objective and Scope

The first step in constructing an index number is to define its purpose clearly. The objective may be to measure changes in prices, quantities, or values over time or between regions. This determines whether a price index, quantity index, or value index is required. Additionally, the scope must be outlined—whether it’s for a particular sector (like retail or wholesale prices) or a specific group (such as urban consumers). Defining the objective ensures relevance, appropriate selection of items, and accurate interpretation of the index in practical use.

2. Selection of the Base Year

The base year is the reference year against which changes are compared. It is assigned a value of 100, and all subsequent values are calculated in relation to it. The base year should be a “normal” year—free from major economic disruptions like inflation, war, or natural disasters. A poorly chosen base year may distort the index. Additionally, it should be recent enough to reflect current trends but stable enough to serve as a benchmark. Periodic updating of the base year is essential for long-term accuracy.

3. Selection of Commodities

Next, a representative basket of goods and services must be selected. These commodities should reflect the consumption habits or production patterns of the population or sector under study. Items should be commonly used, available throughout the period, and consistent in quality. Too many items can complicate calculations, while too few may result in an unrepresentative index. For example, the Consumer Price Index includes food, clothing, fuel, and transportation. Proper selection ensures the index accurately reflects real economic conditions and consumer behavior.

4. Collection of Price Data

Prices for the selected commodities must be collected for both the base year and the current year. This data should be gathered from reliable sources such as retail shops, wholesale markets, or government reports. Consistency in quality, unit, and location is crucial to ensure accuracy. Prices may vary by region, seller, or time, so care must be taken to eliminate anomalies. Regular and systematic price collection—monthly or quarterly—is often used in official indices. Errors or inconsistencies in this stage can significantly affect the results.

5. Assigning Weights

Weights represent the relative importance of each commodity in the index. Heavier weights are given to items with a larger share in total expenditure or production. For instance, in a household index, food items may carry more weight than luxury goods. Assigning correct weights helps the index reflect real economic behavior. Weights can be based on surveys, national accounts, or expenditure studies. There are unweighted indices (equal importance to all items) and weighted indices (varying importance), with weighted indices offering greater precision and realism.

6. Selection of the Index Formula

Different formulas are used to calculate the index number. The most common are:

  • Laspeyres’ Index: Uses base year quantities as weights.

  • Paasche’s Index: Uses current year quantities.

  • Fisher’s Ideal Index: Geometric mean of Laspeyres and Paasche indices.

Each formula has its pros and cons. Laspeyres is easier to calculate but may overstate inflation, while Paasche may understate it. Fisher’s index balances both but is more complex. The choice depends on available data and desired accuracy. The selected formula must ensure consistency and logical interpretation.

7. Computation and Interpretation

Once the prices, quantities, weights, and formula are determined, the index number is computed. The resulting figure indicates the level of change compared to the base year. If the index is above 100, it shows a price rise; below 100 indicates a fall. The index is then interpreted in the context of economic conditions and published for use by policymakers, businesses, and researchers. Proper interpretation helps in understanding inflation trends, making wage adjustments, or planning fiscal and monetary policies effectively.

Simple Average or Price Relative Method, Weighted index method

Simple Average or Price Relatives Method

In this method, we find out the price relative of individual items and average out the individual values. Price relative refers to the percentage ratio of the value of a variable in the current year to its value in the year chosen as the base.

Price relative (R) = (P1÷P2) × 100

Here, P1= Current year value of item with respect to the variable and P2= Base year value of the item with respect to the variable. Effectively, the formula for index number according to this method is:

 P = ∑[(P1÷P2) × 100] ÷N

Here, N= Number of goods and P= Index number.

Weighted index method

Weighted Aggregate Method

Here different goods are assigned weight according to the quantity bought. There are three well-known sub-methods based on the different views of economists as mentioned below:

Laspeyre’s Method

Laspeyre was of the view that base year quantities must be chosen as weights. Therefore the formula is :

P = (∑P1Q0÷∑P0Q0)×100

Here,  ∑P1Q0= Summation of prices of current year multiplied by quantities of the base year taken as weights and ∑P0Q0= Summation of, prices of base year multiplied by quantities of the base year taken as weights.

Paasche Index Number

The Paasche Price Index is a consumer price index used to measure the change in the price and quantity of a basket of goods and services relative to a base year price and observation year quantity. Developed by German economist Hermann Paasche, the Paasche Price Index is commonly referred to as the “current weighted index.”

Formula for the Paasche Price Index

The formula for the index is as follows:

Where:

  • Pi,0 is the price of the individual item at the base period and Pi,t is the price of the individual item at the observation period.
  • Qi,t is the quantity of the individual item at the observation period.

Marshall Edgeworth Index Number

Tests of Adequacy (TRT and FRT)

To ensure the reliability and accuracy of an index number, it must satisfy certain mathematical tests of consistency, known as Tests of Adequacy. The two most important tests are:

Time Reversal Test (TRT):

Time Reversal Test checks the consistency of an index number when time periods are reversed. In other words, if we calculate an index number from year 0 to year 1, and then from year 1 back to year 0, the product of the two indices should be equal to 1 (or 10000 when expressed as percentages).

Mathematical Condition:

P01 × P10 = 1

or

P01 × P10 = 10000

Where:

  • P01 = Price index from base year 0 to current year 1

  • P10 = Price index from current year 1 to base year 0

Interpretation:

This test ensures that the index number gives symmetrical results when the time order of comparison is reversed.

Which Formula Satisfies TRT?

  • Fisher’s Ideal Index satisfies the Time Reversal Test.

  • Laspeyres’ and Paasche’s indices do not satisfy this test.

Factor Reversal Test (FRT):

Factor Reversal Test checks whether the product of the Price Index and the Quantity Index equals the value ratio (i.e., the ratio of total expenditure in the current year to that in the base year).

Mathematical Condition:

P01 × Q01 = ∑P1Q1 / ∑P0Q0

Where:

  • P01 = Price index from base year to current year

  • Q01 = Quantity index from base year to current year

  • ∑P1Q1 = Total value in the current year

  • ∑P0Q0 = Total value in the base year

Interpretation:

This test checks whether the index number captures the combined effect of both price and quantity changes on total value.

Which Formula Satisfies FRT?

  • Fisher’s Ideal Index satisfies the Factor Reversal Test.

  • Laspeyres’ and Paasche’s indices do not satisfy this test.

Consumer Price Index

Consumer Price Index is also known as the cost of living index.

It represents the average change in price over a period of time, paid by a consumer for a fixed basket of goods and services.

Uses of CPI:

  • It indicates the changes in the consumer prices.
  • It evaluates the purchasing power of money.
  • It is also used for comparison purposes.

Limitations of CPI;

  • CPI focuses on a fixed basket, as consumer behaviour cannot be predicted, we can’t be very sure about CPI value to be relevant.
  • Quality is not considered while calculating the CPI.
  • Inflation effects are not taken into consideration as the basket is fixed.

CPI can be computed using 2 methods:

  • Aggregate Expenditure method

CPI = (Total expenditure in current year/Total expenditure in base year)*100; which means;

CPI = Σp1q0/Σp0q0 * 100

  • Family Budget method

CPI =  ΣWP/ ΣW

Where P = p1/p0 * 100

Smoothed frequency curve

The frequency is the number of times an event occurs within a given scenario. Cumulative frequency is defined as the running total of frequencies. It is the sum of all the previous frequencies up to the current point. It is easily understandable through a Cumulative Frequency Table.

Marks Frequency

(No. of Students)

Cumulative Frequency
0 – 5 2 2
5 – 10 10 12
10 – 15 5 17
15 – 20 5 22

Cumulative Frequency is an important tool in Statistics to tabulate data in an organized manner. Whenever you wish to find out the popularity of a certain type of data, or the likelihood that a given event will fall within certain frequency distribution, a cumulative frequency table can be most useful. Say, for example, the Census department has collected data and wants to find out all residents in the city aged below 45. In this given case, a cumulative frequency table will be helpful.

Cumulative Frequency Curve

A curve that represents the cumulative frequency distribution of grouped data on a graph is called a Cumulative Frequency Curve or an Ogive. Representing cumulative frequency data on a graph is the most efficient way to understand the data and derive results.

There are two types of Cumulative Frequency Curves (or Ogives) :

  • More than type Cumulative Frequency Curve
  • Less than type Cumulative Frequency Curve

Frequency Polygon

A frequency polygon is a graphical form of representation of data. It is used to depict the shape of the data and to depict trends. It is usually drawn with the help of a histogram but can be drawn without it as well. A histogram is a series of rectangular bars with no space between them and is used to represent frequency distributions.

Steps to Draw a Frequency Polygon

  • Mark the class intervals for each class on the horizontal axis. We will plot the frequency on the vertical axis.
  • Calculate the classmark for each class interval. The formula for class mark is:

Classmark = (Upper limit + Lower limit) / 2

  • Mark all the class marks on the horizontal axis. It is also known as the mid-value of every class.
  • Corresponding to each class mark, plot the frequency as given to you. The height always depicts the frequency. Make sure that the frequency is plotted against the class mark and not the upper or lower limit of any class.
  • Join all the plotted points using a line segment. The curve obtained will be kinked.
  • This resulting curve is called the frequency polygon.

Note that the above method is used to draw a frequency polygon without drawing a histogram. You can also draw a histogram first by drawing rectangular bars against the given class intervals. After this, you must join the midpoints of the bars to obtain the frequency polygon. Remember that the bars will have no spaces between them in a histogram.

Question 1: Construct a frequency polygon using the data given below:

Test Scores Frequency
49.5-59.5 5
59.5-69.5 10
69.5-79.5 30
79.5-89.5 40
89.5-99.5 15

Answer: We first need to calculate the cumulate frequency from the frequency given.

Test Scores Frequency Cumulative Frequency
49.5-59.5 5 5
59.5-69.5 10 15
69.5-79.5 30 45
79.5-89.5 40 85
89.5-99.5 15 100

We now start by plotting the class marks such as 54.5, 64.5, 74.5 and so on till 94.5. Note that we will also plot the previous and next class marks to start and end the polygon, i.e. we plot 44.5 and 104.5 as well.

Then, the frequencies corresponding to the class marks are plotted against each class mark. Like you can see below, this makes sense as the frequency for class marks 44.5 and 104.5 are zero and touching the x-axis. These plot points are used only to give a closed shape to the polygon. The polygon looks like this:

error: Content is protected !!