This opens the Data Science & AI track. It argues for an unpopular starting point: not a model, not a framework, but about six statistical ideas.
The reason is practical. People who start with models can run a training script in week two and cannot tell whether the result means anything in month six. People who start with the statistics are slower for a fortnight and then correct.
The six ideas
1. Distributions
Before any analysis, look at the shape of your data. Is it symmetric, skewed, bimodal? Does it have a long tail? A mean reported on a skewed distribution is misleading, and most real-world data โ incomes, waiting times, file sizes, page views โ is skewed.
The habit: plot a histogram of every variable before you do anything else. It takes one line and prevents entire categories of mistake.
2. Variance and standard deviation
Two datasets can share a mean and have nothing else in common. Spread tells you how much to trust a single number. Reporting an average without a measure of spread is the most common way analyses mislead.
3. Correlation is not causation, and correlation is also not the only relationship
The first half is famous. The second half matters more in practice: a correlation of zero does not mean no relationship, it means no linear relationship. A perfect U-shape has near-zero correlation. Again: plot it.
4. Sampling and bias
Your data came from somewhere, and where it came from determines what it can tell you. A survey of app users tells you about people who kept the app, not about people who deleted it. This is the source of more wrong conclusions than any modelling error.
Ask, every time: who or what is missing from this dataset?
5. Statistical significance and what it does not mean
A p-value is the probability of seeing a result at least this extreme if nothing were going on. It is not the probability that your hypothesis is true, and it says nothing about whether the effect is large enough to care about. With a big enough sample, meaningless differences become significant.
Always report the effect size next to the significance.
6. Overfitting
A model that fits your existing data perfectly has probably memorised its noise. This is why you hold data back and test on it. Understand this before you touch a model and you will avoid the most common beginner result: 99 percent accuracy in training and useless in practice.
A sensible first project
Take a dataset you care about โ your own grades, your spending, a public dataset on a topic you find interesting โ and do this without any machine learning:
- Load it and count the rows. Then count the missing values per column.
- Plot the distribution of every numeric column.
- Write down three questions you want answered.
- Answer them with group-bys and simple summaries.
- Write a paragraph on what the data cannot tell you and why.
That last step is the one that develops judgement, and judgement is what distinguishes a data scientist from someone who can call a library.
The tooling, briefly
Python with a dataframe library and a plotting library is the standard route and there is no reason to deviate. Learn to load, filter, group, join and plot. Those five operations cover the large majority of real analysis work.
On using AI tools while learning this
They are good at writing the code and poor at knowing whether the question was sensible. Use them for syntax you have forgotten, and do the thinking about what to measure yourself โ that thinking is the job.
Next in this track
Part 2 covers cleaning messy data, which is where most of the time in real work goes. Part 3 covers your first model and how to evaluate it honestly.