← Back to list

Parsing Raw CAN Logs Without Losing Your Mind

Raw CAN logs are not friendly artifacts. They are not formatted for humans.

Aeon Flex, Elriel Assoc. 2133 [NEON MAXIMA] in Coding Nexus · 2026-03-03 16:05 · 1 claps · 5.3 min read
#can-bus #algorithms #hardware-hacking #car-hacking #showdev
Open on Medium ↗
Wiki topics: 💻 · Programming 🔒 · Cybersecurity 🏺 · Archaeology & Anthropology

Parsing Raw CAN Logs Without Losing Your Mind

Photo by Nikhil . on Unsplash

Photo by Nikhil . on Unsplash

It is 2:17 a.m. The garage smells like burnt insulation and stale coffee. Your flashlight rolls across concrete and stops under the chassis. The OBD-II connector is half exposed, wires like nerves pulled from a spine. The laptop hums. Terminal open. Candump scrolling. Hex flooding the screen faster than your eyes can metabolize it.

This is the ritual.

Raw CAN logs are not friendly artifacts. They are not formatted for humans. They are machine whispers timestamped in microseconds, arbitration IDs colliding in disciplined silence, payloads packed into eight bytes like contraband. You do not read them. You endure them. Then, if you are patient, you begin to understand them.

At first glance a log looks like someone liquified ten spreadsheets and poured them into a text file.

(158264.123456) 1F4#0A1B2C3D4E5F6789

Timestamp. ID. Payload. That is it. No legend. No DBC. No apology.

You will not decode this by staring at it harder. You will not unlock it by sheer will. What you need is orientation.

Accept That It Looks Insane

The first mistake people make is expecting clarity.

A raw CAN stream is a continuous river of frames. Each frame contains an arbitration ID and up to eight bytes of data. Sometimes less. Sometimes padded. Sometimes fragmented across multiple frames if ISO-TP is involved. If you are lucky, you know the bus speed and whether it is standard or extended ID. If you are unlucky, you are guessing.

You look at:

1F4#0A1B2C3D4E5F6789

And you see nothing. Good. That means you are honest.

The structure is simple. The meaning is not. Meaning hides in repetition. In delta. In timing. In absence.

Your job is not to understand everything at once. Your job is to build a mental map of the battlefield.

Count Before You Decode

Before you speculate about what byte 3 represents, count.

Which IDs appear most frequently. Which IDs burst in clusters. Which ones appear only when you press the brake or toggle ignition. Extract frequency distributions. Calculate inter-frame timing deltas. Build histograms. You are not decoding yet. You are profiling.

A minimal Python parser is enough to begin. Read the file line by line. Split timestamp from ID and payload. Track occurrence counts per arbitration ID. Track average delta time per ID. Dump to CSV. Plot.

You are building terrain awareness.

This is also where a tool I just finished over the weekend, CANBUSconfidenceid, enters like a disciplined scout instead of a hype machine. It does not pretend to decode the bus for you. It performs structural inference. It hunts for rolling counters. It evaluates checksum hypotheses. It scores entropy across each byte position.

Point it at a raw candump log and it returns probabilistic statements. Not vibes. Not guesses.

Byte 6 of ID 0x1F4 matches XOR of bytes 0 through 5 in 92 percent of frames.

Byte 0 increments monotonically modulo 16.

Byte 3 shows high entropy consistent with sensor scaling.

That collapses hours of blind correlation into a ranked hypothesis list. You now know where to look.

Run it early.

can-hypothesis input.log — report

Save structured output.

can-hypothesis input.log -o results.json

Do not skip this step because you think you can eyeball it. You cannot. Nobody can at scale.

See the Delta

Hex in a terminal is static. Machines change. You need motion.

This is where visualization stops being optional. It becomes survival.

The other tool I released over the weekend- CAN Frame Playground- was built for exactly this stage of fatigue. Drag in a candump file. Instantly you see syntax highlighting. More importantly, per byte delta detection. Bytes that changed from the previous frame glow. Static bytes fade into background noise.

Your eyes are wired to detect change. Let the interface serve that biology.

The delta graph tab lets you select an arbitration ID and plot byte changes over time. Suddenly what looked like 0A 1B 2C 3D becomes a waveform. Spikes correlate to pedal presses. Plateaus correlate to idle states. Flatlines become background processes.

The annotation layer matters more than people think. You confirm that byte 0 is a rolling counter. Label it. Byte 3 scales linearly with vehicle speed. Label it. Over hours the log transforms from entropy into documentation.

All local processing. No upload. No cloud. No accidental leakage of data that was never meant to be public.

If you prefer CLI asceticism, fine. Export to CSV. Plot with matplotlib. The principle does not change. Visualize change. Investigate spikes. Ignore static.

Decode One ID at a Time

Nobody decodes an entire vehicle network in one pass. That is fantasy.

Pick one arbitration ID. Ideally one flagged by CANBUSconfidenceid as containing a rolling counter or checksum. Anchors reduce ambiguity. If byte 0 is confirmed as a counter, you subtract that from the unknown pool. One variable eliminated.

Track payload changes frame by frame.

Example translation table:

0A -> 10 km/h

14 -> 20 km/h

1E -> 30 km/h

You do not guess scaling immediately. You record correlations. Drive at steady 20 km/h. Observe payload. Repeat at 30. Confirm linearity. Document.

Once you decode one signal inside a frame, the rest of the payload becomes easier to isolate because you have contextual grounding. Speed correlates with RPM. RPM correlates with throttle position. Correlation is leverage.

Do not chase five IDs simultaneously. That is how cognitive drift happens. Focus. Extract. Annotate. Move outward.

Entropy Is a Map

Entropy analysis is not academic decoration. It is structural reconnaissance.

Low entropy byte positions across a session often indicate flags or mode bits. High entropy positions likely contain live sensor data. Moderate entropy with modular patterns often indicates rolling counters or scaled integers.

Run entropy profiling per ID. Plot it. You will see which bytes deserve attention.

Your brain loves narrative. Entropy gives it scaffolding.

Document Like a Prosecutor

Every assumption goes into notes. Every mapping. Every anomaly. Timestamp it. Add physical context. Ignition state. Vehicle speed. Button pressed. Engine warm or cold.

Treat it like a case file. If you cannot return in six months and understand your own reasoning, you failed documentation discipline.

CAN Playground handles inline annotation. Supplement it with a markdown file. Write why you believed byte 4 was brake pressure. Write why you abandoned that hypothesis.

Future you will thank you. Or curse you if you do not.

Sanity Protocol

There is a psychological component nobody talks about.

After three hours of hex, your brain starts hallucinating patterns. You will see structure where none exists. This is normal. It is also dangerous.

Step away. Touch hardware. Look at the physical connector. Feel the weight of the harness. Remind yourself there is a machine behind the data.

Small wins matter. Decoding one reliable ID is progress. Identifying a checksum is progress. Each confirmed hypothesis reduces chaos.

Do not attempt full automation before orientation. Automation without understanding produces confident nonsense.

Scaling the Method

Once you have a repeatable workflow, larger logs stop being overwhelming.

Multiple capture sessions. Firmware comparisons. Different vehicles of the same model. The same pipeline scales.

Frequency analysis. Hypothesis scoring. Visualization. Incremental decode. Documentation.

Patterns repeat across architectures. Rolling counters look similar. Checksums follow predictable families. Scaling becomes multiplication of discipline, not reinvention.

If you move into distributed parsing across nodes or lightweight clusters, the mindset remains identical. Parallelization does not replace comprehension. It amplifies it.

The Real Reward

CAN frames were never meant for you. They are machine to machine communication optimized for arbitration latency and fault tolerance, not legibility.

Decoding them is not just reverse engineering. It is translation.

The first time a raw candump stops looking like noise and starts reading like a language, something shifts. You are no longer reacting to hex. You are interpreting state transitions. You are predicting behavior.

That fluency compounds.

Every ID you name weakens opacity. Every checksum you identify strips away assumed secrecy. The bus becomes less of a black box and more of a conversation.

And at 3:40 a.m., when the garage is quiet and the terminal is no longer hostile but familiar, you realize the goal was never just decoding.

It was learning how to see.


메타데이터
post_id
4f735e1c74b1
slug
parsing-raw-can-logs-without-losing-your-mind-4f735e1c74b1
url
https://medium.com/@neonmaxima/parsing-raw-can-logs-without-losing-your-mind-4f735e1c74b1
canonical_url
https://medium.com/@neonmaxima/parsing-raw-can-logs-without-losing-your-mind-4f735e1c74b1
author_url
https://medium.com/@neonmaxima
status
ok
fetched_at
2026-06-22 05:41:33