Ceph's thresholds for what deserves going into CEPH_WARN state is always curious. For example: Like, ok, cool So I go and look at the sled that has that OSD on it, and run "dmesg" as a general vibe check. Cool, /dev/sdh is the drive that seems to be flapping every few days. Actually, nope! Ceph had picked up that a totally different drive had returned 3 read errors in 24 hours, that triggered CEPH_WARN. But a totally different OSD just falling off, and coming back for a week+ (and the SMART has a lot of (successful) reassigned sectors), not worth even slightly talking about. Curious software
BLUESTORE_SPURIOUS_READ_ERRORS: 1 OSD(s) have spurious read errors
BarbarossaTM@noc.soc..
replied 15 Aug 2026 14:49 +0000
in reply to: https://benjojo.co.uk/u/benjojo/h/6kWGTQ8LvS14779pXs
benjojo
replied 15 Aug 2026 15:14 +0000
in reply to: https://noc.social/users/BarbarossaTM/statuses/117100076447917587
@BarbarossaTM 16 😬, since it's what debian 12 ships with, pending a upgrade very soon to 18, and then hopefully upwards to 20
benjojo
replied 15 Aug 2026 15:45 +0000
in reply to: https://noc.social/users/BarbarossaTM/statuses/117100206874389819
vjon@mastodon.online
replied 16 Aug 2026 19:53 +0000
in reply to: https://benjojo.co.uk/u/benjojo/h/6kWGTQ8LvS14779pXs
@benjojo Ceph has a place in my heart. Right next to Reiser4. The only filesystems which ate my data without any hardware failure as a trigger. And yeah, the Ceph logs about the issue were less than useful.
benjojo
replied 17 Aug 2026 10:57 +0000
in reply to: https://mastodon.online/users/vjon/statuses/117106933216074633
@vjon never used Reiser4, but I can believe that ceph if held in a way can lose data, it's immensely complicated and unforgiving, unfortunately it also becomes one of the view viable things once you cross a volume of storage...