Konrad Kowalski (rootsher)Principal Platform & Reliability Architect111111101111000011010100110011100111010111000111

Performance Tests and Load Tests: Behavior Under Traffic

date
category
Testing
also in
Capacity & Performance · Observability
reading
2 min / 490 words

A system can return the correct answer and still be broken.

If checkout works for one user but responds after five seconds under normal traffic, correctness testing did not save anything.

That is where performance tests, load tests, and stress tests belong.

What kind of test this is

Performance test measures system behavior over time.

It can check:

text
latency
throughput
startup time
CPU usage
memory usage
number of database queries
rendering time

Load test checks the system under expected traffic.

Stress test intentionally goes beyond expected traffic to find the boundary.

These names are close, but they do not mean the same thing.

Performance test

A performance test can be small.

It does not always mean a large traffic test.

Examples:

text
a function processes 10,000 records below 200 ms
an endpoint does not make more than 3 database queries
table rendering fits the time budget
a worker processes a batch without memory growth

This kind of test catches cost regression.

The code still gives the correct result, but does it more slowly or more expensively.

Load test

A load test asks:

text
can the system handle the traffic we expect?

That requires a scenario.

A number of requests per second is not enough.

We need to know:

text
which endpoints
traffic distribution
how many users
how many writes
how many reads
what data
how long

A load test without a realistic traffic profile is easy to run and hard to interpret.

Stress test

A stress test asks:

text
where does the system break?

This is not a daily gate for every pull request.

It is a tool for learning boundaries.

A good stress test says not only at which load the system stops meeting requirements.

It also says how it stops working.

text
does latency grow
do 500 errors appear
does the database hit CPU
does the queue grow
does autoscaling keep up
does the system degrade gracefully

What happens when they are missing

Without performance tests, performance remains an opinion.

Someone says "it should be enough".

Someone looks at average response time.

Someone else checks only locally.

The problem appears only when traffic is real.

Then production becomes the test.

If there is no observability, even the result of that test is hard to read.

Budgets instead of vague expectations

A performance test should have a budget.

text
p95 below 300 ms
fewer than 5 SQL queries
startup below 2 s
batch below 1 GB RAM
checkout handles 200 RPS for 15 minutes

Without a budget, the test only measures.

With a budget, it can stop a regression.

The budget does not have to be perfect from the start.

It is better to have a simple boundary and adjust it than to have charts without a decision.

Tool from the stack

For load tests I would choose k6.

It is simple enough to describe the scenario in code and close enough to DevOps to run from the pipeline or manually before a larger change.

For a Node backend I would use it against public endpoints and critical API flows.

For the frontend it does not replace rendering tests, but it can check the backend cost of paths used by React.

A useful target looks like this:

text
checkout API keeps p95 below 300 ms at expected RPS

Without a budget, k6 is only a traffic generator.

With a budget, it becomes a quality gate.

When the test misleads

A performance test misleads when the environment does not resemble the problem.

text
different database
different indexes
different CPU limits
no TLS
no cache miss
too small dataset
no write contention

Not every test has to be production.

But we need to know what the test does not simulate.

Otherwise the result looks scientific but describes a different system.

The final part of the series covers a layer often called configuration, even though it can break the application as effectively as code.