Did the Linux boot really fail? Read a boot, the previous boot and an unexpected reboot

The short answer: when you read Linux boot logs, a Linux boot is not “failed” because the log has red lines. It failed if systemctl is-system-running and systemctl --failed say so. If the machine rebooted by itself, the evidence is in the end of the previous boot, and you can only read that if the journal is persistent.

How to read this article. Every log line and output carries a label. Lab output means the author ran it on a Rocky Linux 8.10 virtual machine (VMware) and pasted the real result; names were replaced with host01 and long lines were trimmed where marked. Documented means it comes from official documentation, with a link. Anything the author did not run is explained from the official documentation and marked Documented.

What it is

Term Plain meaning
Boot Everything from power-on to a usable system: firmware, bootloader, kernel, initramfs, systemd.
systemd The first user-space process (PID 1). It starts services (called units) and also writes logs.
journald and the journal systemd-journald collects kernel messages and service output into a binary log called the journal. You read it with journalctl.
Volatile journal Stored in /run/log/journal, which lives in memory. Lost at reboot.
Persistent journal Stored in /var/log/journal on disk. Survives reboot.
Boot ID A random ID made at each boot. The journal tags every line with it, so lines can be grouped per boot.
Kernel ring buffer A small in-memory log the kernel writes to. dmesg reads it. It holds the current boot only.

Do not confuse the systemd journal with the filesystem journal of ext4 or XFS. They are different things. The filesystem journal only matters here as evidence: if a filesystem has to replay its journal at boot, the previous stop was not clean.

Why it matters

Mistake Real impact
Treating harmless warnings as the cause Hours lost chasing noise while the real fault (a failed unit, a full disk) is missed.
Volatile journal and an unexpected reboot The only log of what happened is gone. You cannot tell a power cut from a kernel panic.
Not investigating an unexpected reboot It repeats. On a storage or backup server it can also mean unsynced writes, filesystem recovery and backup jobs that stopped halfway.
No visible console on a hung boot You reset the machine blindly and lose the evidence.

How it works

power on -> firmware -> bootloader (GRUB) -> kernel -> initramfs -> systemd (PID 1) -> units -> target reached
                                               |                        |
                                       kernel ring buffer        journald collects kernel + units
                                       (dmesg, RAM only)                |
                                               +--------> journal <-----+
                                   volatile:   /run/log/journal   (lost at reboot)
                                   persistent: /var/log/journal   (kept)

The Storage= setting in /etc/systemd/journald.conf decides where the journal goes.

Storage= Behaviour
volatile Only /run/log/journal. Lost at reboot.
persistent /var/log/journal, created if missing.
auto Persistent only if /var/log/journal already exists, otherwise volatile.
none Nothing is stored.

Documented: journald.conf(5) describes these four behaviours and the drop-in directory /etc/systemd/journald.conf.d/. Whether /var/log/journal exists after install depends on the distro and version, so check your own machine.

journalctl -b 0 is the current boot, -b -1 the one before, -b -2 the one before that. Those offsets only work if the older boots were stored.

Where you see it: the Linux boot logs and evidence

Each command below shows where the evidence lives and whether it survives a reboot.

journalctl -b                     # journal, current boot (survives reboot only if the journal is persistent)
journalctl -b -1                  # journal, previous boot (only if persistent)
journalctl --list-boots           # list of stored boots
dmesg -T                          # kernel messages, current boot only (RAM, lost at reboot)
journalctl -k                     # kernel messages from the journal (older boots too, if persistent)
last -x reboot shutdown           # boot and shutdown records from /var/log/wtmp (survives reboot)
systemd-analyze time              # boot timing (also: blame, critical-chain), computed for the current boot
systemctl --failed                # failed units, current boot
systemctl is-system-running       # overall state, current boot
cat /proc/cmdline                 # kernel parameters used for this boot

Outside the operating system, the BMC or hypervisor event log also survives a reboot. For power or hardware causes it is often the only evidence.

Case A: a noisy boot that is not a failure

Start with the result, not the noise.

systemctl is-system-running
systemctl --failed

Lab output (Rocky Linux 8.10, kernel 4.18, VMware VM):

degraded
  UNIT                   LOAD   ACTIVE SUB    DESCRIPTION
● pmlogger_daily.service loaded failed failed Process archive logs

Documented: the is-system-running states are initializing, starting, running, degraded, maintenance, stopping, offline and unknown; degraded means one or more units failed (systemctl(1)).

What this tells you: the state is degraded because exactly one unit failed, pmlogger_daily.service (“Process archive logs”). The author did not investigate why it failed. The machine was fully usable. So degraded means “one unit failed, go and look at it”, not “the machine is broken”. The next step is to look at that unit:

systemctl status pmlogger_daily.service
journalctl -b -u pmlogger_daily.service

Now the error-level lines of the same boot:

journalctl -b -p err --no-pager

Lab output (five of the lines, hostname replaced, others omitted):

Sep 27 12:19:35 host01 kernel: piix4_smbus 0000:00:07.3: SMBus Host Controller not enabled!
Sep 27 12:19:38 host01 kernel: Bluetooth: hci0: unexpected cc 0x0c12 length: 2 < 3
Sep 27 12:19:38 host01 kernel: Bluetooth: hci0: Opcode 0x c12 failed: -38
Sep 27 12:19:42 host01 alsactl[1050]: alsa-lib main.c:1554:(snd_use_case_mgr_open) error: failed to import hw:0 use case configuration -2
Sep 27 12:20:07 host01 libvirtd[1226]: Unable to open /dev/kvm: No such file or directory
Line What it tells you
piix4_smbus ... SMBus Host Controller not enabled! The virtual machine has no SMBus controller to talk to. Harmless on a VM.
Bluetooth: hci0: ... failed: -38 A virtual Bluetooth device that does not support a command. Harmless here.
alsactl ... failed to import hw:0 use case configuration Audio setup on a machine with no real sound hardware. Harmless here.
libvirtd ... Unable to open /dev/kvm No hardware virtualisation inside this VM, so KVM is not available. Expected when running inside a VM.

These are all priority err, so they look serious. None of them stopped the machine from booting. The rule: judge a boot by systemctl --failed, then use the log to explain the failures, not to find them. Keep a short list of “known harmless” lines per platform so the team stops chasing them.

Case B: a slow boot is not a failed boot

systemd-analyze time
systemd-analyze blame
systemd-analyze critical-chain

Lab output:

Startup finished in 4.585s (kernel) + 9.047s (initrd) + 1min 55.465s (userspace) = 2min 9.098s
graphical.target reached after 1min 55.413s in userspace
    1min 34.888s kdump.service
         39.715s plymouth-quit-wait.service
         21.417s tuned.service
         18.852s unbound-anchor.service
         15.830s pmlogger.service

Lab output (critical chain, trimmed after sockets.target):

graphical.target @1min 55.413s
└─multi-user.target @1min 55.412s
  └─plymouth-quit-wait.service @21.160s +39.715s
    └─systemd-user-sessions.service @20.624s +344ms
      └─remote-fs.target @20.479s
        └─remote-fs-pre.target @20.476s
          └─nfs-client.target @20.065s
            └─gssproxy.service @19.581s +410ms
              └─network.target @19.528s
                └─wpa_supplicant.service @50.874s +464ms
                  └─dbus.service @14.216s
                    └─basic.target @13.409s
                      └─sockets.target @13.409s

What this tells you: the boot took 2 min 9 s, almost all of it in user space. blame lists kdump.service first at 1 min 34 s, but blame only sorts by how long each unit took. Units start in parallel, so a slow unit is not necessarily the one that delayed the boot. critical-chain shows the chain of units that decided when graphical.target was reached. Fix only units on that chain. The boot still finished, so this is a performance question, not a failure.

Case C: the failure was in the previous boot (the journal was volatile)

The outage happened, the machine came back, and the current boot looks clean. The next question is whether the previous boot was stored.

ls -ld /var/log/journal
ls -ld /run/log/journal
journalctl --list-boots
journalctl -b -1 -n 30 --no-pager

Lab output (before any change on this VM):

ls: cannot access '/var/log/journal': No such file or directory
drwxr-sr-x. 3 root systemd-journal 60 Sep 27 12:19 /run/log/journal
 0 59c55804affc4ccfb30297b82ec7ed53 Sun 2026-09-27 12:19:19—Sun 2026-10-04 12:33:20
Specifying boot ID or boot offset has no effect, no persistent journal was found.

What this tells you:

  • /var/log/journal does not exist and /run/log/journal does, so the journal is volatile.
  • --list-boots shows one boot only. This VM had been up for seven days, and all that time was one boot.
  • journalctl -b -1 fails with a clear message: no persistent journal was found.

If this had been an outage, the log of the previous boot would be gone. The remaining evidence would be last -x, text logs (if rsyslog is running), and the hypervisor or BMC event log.

Fix: make the journal persistent

THIS CHANGES THE SYSTEM. Disposable VM first, with a snapshot and a console you can reach outside SSH. On a real server, use your change process. Rollback: delete the drop-in file below and run sudo systemctl restart systemd-journald.
sudo mkdir -p /etc/systemd/journald.conf.d
printf '[Journal]\nStorage=persistent\n' | sudo tee /etc/systemd/journald.conf.d/10-persistent.conf
sudo systemctl restart systemd-journald
sudo journalctl --flush

Documented: journalctl --flush moves data from /run/log/journal/ to /var/log/journal/ when persistent storage is enabled (journalctl(1)). Journals have a size cap: SystemMaxUse= defaults to 10 percent of the filesystem, capped at 4G (journald.conf(5)). Check yours with journalctl --disk-usage.

Then reboot, and check:

journalctl --list-boots

Lab output (after the first reboot):

-1 59c55804affc4ccfb30297b82ec7ed53 Sun 2026-09-27 12:19:19—Sun 2026-10-04 12:48:21
 0 b474385f8f2f4280b1f63a355337d951 Sun 2026-10-04 12:48:40—Sun 2026-10-04 12:51:21

Two boots are listed now: the old one (-1) and the current one (0). That is the proof that the fix worked.

Case D: an unexpected reboot, clean or unclean?

Read the end of the previous boot. The jump to the end is the -e option.

journalctl -b -1 -e --no-pager
last -x reboot shutdown -n 5

Clean stop, Lab output (a planned sudo reboot; last lines of the previous boot):

Oct 04 12:53:04 host01 systemd[1]: Reached target Reboot.
Oct 04 12:53:04 host01 systemd[1]: Shutting down.
Oct 04 12:53:05 host01 kernel: printk: systemd-shutdow: 38 output lines suppressed due to ratelimiting
Oct 04 12:53:05 host01 systemd-shutdown[1]: Syncing filesystems and block devices.
Oct 04 12:53:05 host01 systemd-shutdown[1]: Sending SIGTERM to remaining processes...
Oct 04 12:53:05 host01 systemd-journald[696]: Journal stopped
reboot   system boot  4.18.0-553.168.1 Sun Oct  4 12:53   still running
shutdown system down  4.18.0-553.168.1 Sun Oct  4 12:53 - 12:53  (00:00)
reboot   system boot  4.18.0-553.168.1 Sun Oct  4 12:48 - 12:53  (00:04)
shutdown system down  4.18.0-553.168.1 Sun Oct  4 12:48 - 12:48  (00:00)
reboot   system boot  4.18.0-553.168.1 Sun Sep 27 12:19 - 12:48 (7+00:28)

A clean stop leaves a trail. The tail has Shutting down, Syncing filesystems and block devices and Journal stopped, and last -x shows a shutdown record before each reboot.

Lab output (a little earlier in the first clean shutdown; the filesystem on /boot is unmounted properly):

Oct 04 12:48:17 host01 systemd[1]: Unmounting /boot...
Oct 04 12:48:17 host01 kernel: XFS (nvme0n1p1): Unmounting Filesystem

Unclean stop, Lab output. The VM was reset from the hypervisor instead of being shut down from inside. Last lines of the previous boot (the last three long desktop lines are omitted, and two further lines were removed):

Oct 04 12:55:36 host01 systemd[1]: NetworkManager-dispatcher.service: Succeeded.
Oct 04 12:55:36 host01 realmd[4183]: quitting realmd service after timeout
Oct 04 12:55:36 host01 realmd[4183]: stopping service
Oct 04 12:55:36 host01 systemd[1]: realmd.service: Succeeded.
Oct 04 12:55:44 host01 systemd[1]: systemd-hostnamed.service: Succeeded.
Oct 04 12:55:46 host01 systemd[1]: systemd-localed.service: Succeeded.
Oct 04 12:55:55 host01 kernel: hrtimer: interrupt took 10518827 ns
Oct 04 12:55:56 host01 PackageKit[3994]: resolve transaction /473_dacabbbe from uid 1000 finished with success after 27188ms
Oct 04 12:56:04 host01 systemd[1]: libvirtd.service: Succeeded.
reboot   system boot  4.18.0-553.168.1 Sun Oct  4 12:56   still running
reboot   system boot  4.18.0-553.168.1 Sun Oct  4 12:53   still running
shutdown system down  4.18.0-553.168.1 Sun Oct  4 12:53 - 12:53  (00:00)
reboot   system boot  4.18.0-553.168.1 Sun Oct  4 12:48 - 12:53  (00:04)
shutdown system down  4.18.0-553.168.1 Sun Oct  4 12:48 - 12:48  (00:00)

What this tells you:

  • The tail of the previous boot ends in the middle of ordinary activity. There is no Shutting down and no Journal stopped.
  • last -x shows a reboot record at 12:56 with no shutdown record before it, so there are two reboot ... still running lines in a row.
  • The absence of a clean shutdown record is the clue. It means the stop was unclean: a crash, power loss, a hard reset or a watchdog. The log alone cannot tell these apart. For a hard reset or power loss, the proof is outside the OS, in the hypervisor or BMC event log.

After an unclean stop, also look in the next boot for lines about filesystem recovery. The exact text depends on the filesystem and the kernel version, so search broadly:

journalctl -b 0 --no-pager | grep -iE "recover|uncleanly|not properly"
dmesg | grep -iE "xfs|ext4"

Lab output (the boot after the unclean reset; the first command, hostname replaced):

Oct 04 12:56:42 host01 kernel: XFS (dm-0): Starting recovery (logdev: internal)
Oct 04 12:56:42 host01 kernel: XFS (dm-0): Ending recovery (logdev: internal)
Oct 04 12:57:01 host01 systemd[1]: Starting Crash recovery kernel arming...
Oct 04 12:57:31 host01 systemd[1]: Started Crash recovery kernel arming.

Lab output (the second command; the times in brackets are seconds since the kernel started):

[    9.412036] SGI XFS with ACLs, security attributes, quota, no debug enabled
[    9.424023] XFS (dm-0): Mounting V5 Filesystem
[    9.678375] XFS (dm-0): Starting recovery (logdev: internal)
[    9.918320] XFS (dm-0): Ending recovery (logdev: internal)
[   20.640744] XFS (nvme0n1p1): Mounting V5 Filesystem
[   20.680914] XFS (nvme0n1p1): Ending clean mount

What this tells you:

  • The root filesystem (dm-0) had to replay its journal: Starting recovery, then Ending recovery. XFS only does that when the filesystem was not unmounted cleanly, so this is strong evidence of an unclean stop. It agrees with what the end of the previous boot and last -x already showed.
  • The /boot filesystem (nvme0n1p1) shows Ending clean mount. It needed no recovery.
  • The two Crash recovery kernel arming lines also matched the search, but they are the kdump service starting and have nothing to do with the filesystem. Read each match before drawing a conclusion.

Recovery lines show that the stop was unclean. They do not show why. To tell a crash from a power loss or a hard reset, you still need evidence outside the operating system, such as the hypervisor or BMC event log.

Troubleshooting flow

Work through these in order. Stop as soon as you have your answer.

Step 1. The machine is up and there are many warnings. Check the result first.

systemctl is-system-running
systemctl --failed

running and nothing failed means the warnings are noise. Stop, and do not “fix” them.

Step 1b. The state is degraded, or units are listed as failed. One unit really failed.

systemctl status <unit>
journalctl -b -u <unit>

Follow that unit’s error.

Step 2. The boot is slow. Find which part was slow.

systemd-analyze time
systemd-analyze critical-chain

Fix only units that are on the critical chain.

Step 3. An outage was reported, but the current boot is clean. Is the previous boot stored?

journalctl --list-boots

If yes, go to step 4. If not, the journal was volatile. Use last -x, text logs, and the hypervisor or BMC event log instead.

Step 4. The previous boot exists. Read the end of it, then the errors and the kernel lines.

journalctl -b -1 -e --no-pager
journalctl -b -1 -p err --no-pager
journalctl -b -1 -k --no-pager

Find the first error, not the last.

Step 5. The machine rebooted unexpectedly. Is the stop clean or unclean?

journalctl -b -1 -e --no-pager
last -x reboot shutdown -n 5

A clean stop means someone or something asked for it, so find who. An unclean stop means a crash, power loss or reset, so go to step 6.

Step 6. The stop was unclean. Look for filesystem recovery in the next boot, then look outside the operating system.

journalctl -b 0 --no-pager | grep -iE "recover|uncleanly|not properly"

Also check kdump, pstore and the BMC or hypervisor event log.

The three common causes: harmless warnings mistaken for failure, a volatile journal that hides the past, and a console you are not looking at. The unusual one is a stop that is not software at all, such as power or hardware, where the OS logs contain nothing and the evidence is in the BMC or hypervisor.

If the boot shows nothing on screen

The kernel may be running but printing to another console. Check what it was given:

cat /proc/cmdline

Lab output (Rocky Linux 8.10):

BOOT_IMAGE=(hd0,msdos1)/vmlinuz-4.18.0-553.168.1.el8_10.x86_64 root=/dev/mapper/rl-root ro crashkernel=auto resume=/dev/mapper/rl-swap rd.lvm.lv=rl/root rd.lvm.lv=rl/swap rhgb quiet

There is no console= option here, and quiet and rhgb hide most boot text. To see the messages once, at the GRUB menu press e, remove quiet and rhgb, add console=tty0, and boot with Ctrl-X. The change lasts for one boot only. Documented: GRUB’s menu entry editor boots with Ctrl-X (GNU GRUB manual). Red Hat documents the same steps for RHEL 8: press e at the GRUB menu, find the line that starts with linux, change the parameters on it, press Ctrl-X to boot, and the change applies to that one boot only (Esc cancels the edit) (Red Hat, kernel command-line parameters). Lab output (Rocky Linux 8.10 VM): pressing e at the GRUB menu opens the editor below. The linux line is the one to edit; rhgb quiet are the parameters that hide boot messages. Long lines were cut off by the screen, and the bottom help text is shown as it appeared. After deleting rhgb quiet, adding console=tty0 and booting with Ctrl-X, the running kernel command line shows the edit took effect:

load_video
set gfx_payload=keep
insmod gzio
linux ($root)/vmlinuz-4.18.0-553.168.1.el8_10.x86_64 root=/dev/mapper/rl-root
ro crashkernel=auto resume=/dev/mapper/rl-swap rd.lvm.lv=rl/root rd.lvm.lv=rl/swap rhgb quiet
initrd  ($root)/initramfs-4.18.0-553.168.1.el8_10.x86_64.img $tuned_initrd

Press Ctrl-x to start, Ctrl-c for a command prompt or Escape to discard edits and return to the menu.
cat /proc/cmdline

Lab output (same VM, after the edited boot; quiet and rhgb are gone, console=tty0 is present):

BOOT_IMAGE=(hd0,msdos1)/vmlinuz-4.18.0-553.168.1.el8_10.x86_64 root=/dev/mapper/rl-root ro crashkernel=auto resume=/dev/mapper/rl-swap rd.lvm.lv=rl/root rd.lvm.lv=rl/swap console=tty0

The edit lasted for that boot only; a normal reboot restores the original line. If no new output appears at all, the problem is before the kernel, in GRUB or firmware.

Common mistakes and misconceptions

Mistake Reality
“There are red lines, so the boot failed.” Judge by the result: systemctl --failed.
“journalctl -b -1 always works.” Only with a persistent journal that kept that boot.
“dmesg has the previous boot.” dmesg shows the current boot’s ring buffer only.
“The last line of the log is the cause.” For a crash it is usually just what was happening.
“No shutdown messages means a crash.” It means an unclean stop. Crash, power loss, hard reset or watchdog still have to be told apart.
“The slowest unit in blame is the problem.” Units start in parallel. Look at critical-chain.
“No output on screen means the kernel is dead.” It may be logging to another console.
“I can run fsck to check whether it was clean.” Never run fsck on a mounted filesystem. Read the logs first.

Prevention and monitoring

  • Make the journal persistent on every server and keep a size cap so it cannot fill the disk.
  • Send logs off the machine, so the last lines survive a crash and a lost disk.
  • Alert on uptime reset (“host rebooted”) and on systemctl is-system-running not equal to running.
  • Alert on failed units.
  • Keep a serial or virtual console configured and tested.
  • Learn kdump and pstore, and test them on a VM before relying on them.
  • Write the location of the BMC or hypervisor event log in your runbook.

Quick reference

systemctl is-system-running                  # did the boot really fail?
systemctl --failed                           # which units failed?
journalctl -b -p err --no-pager              # errors in this boot
journalctl --list-boots                      # which boots are stored?
journalctl -b -1 -e --no-pager               # end of the previous boot
last -x reboot shutdown -n 5                 # boots and clean shutdowns
systemd-analyze critical-chain               # why was the boot slow?
ls -ld /var/log/journal                      # is the journal persistent? (then check --list-boots)
cat /proc/cmdline                            # what did the kernel get?

How it was tested

Environment: a Rocky Linux 8.10 virtual machine on VMware, kernel 4.18, used only for this lab. The author ran: the state and failed-unit checks, the error-level journal view, systemd-analyze (time, blame, critical-chain), the volatile-journal checks, making the journal persistent and rebooting, a clean reboot and an unclean reset with journalctl -b -1 and last -x, and cat /proc/cmdline, and the filesystem-recovery search after the unclean stop. The one-boot GRUB edit (remove rhgb quiet, add console=tty0, boot with Ctrl-X) was run and checked with cat /proc/cmdline. Documented only (not run by the author): Ubuntu and Debian, where journal defaults and paths differ by distro. Results may also differ on a systemd version other than the one on this VM.

Sources

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top