This happens after 3-4 days of running the server, then I have to restart it manually.(discuss.tchncs.de)

posted 5 months ago

nutbutter@discuss.tchncs.de

selfhosted@lemmy.world

38 commentshide report

I bought an Optiplex 5040, with an i5-6500TE, and 8 GB DDR3L RAM.

When I bought it, I installed Fedora Server on it. It got stuck every few days but I could never see the error. The services just stopped working, I couldn’t ssh into it, and connecting it to a monitor showed a black screen.

So, I thought let’s install Ubuntu Server, maybe Fedora isn’t compatible with all of its hardware. The same thing is happening, now, but I can see this error. Even when there’s nothing installed on it, no containers, nothing other than base packages, this happens.

I have updated the bios. I have tried setting nouveau.modeset=0 in the grub config file. I have tried disabling and enabling c-states. No luck till now.

Would really appreciate if anyone helps me with this.

UPDATE:

I cleaned everything and reapplied the thermal paste. I did not see any change in the thermals. It never goes over 55°C even under full load.
I reset the motherboard by removing that jumper thing.
I ran memtest86, which took over 2½ hours. It did not show any errors.
I ran a CPU stress test for over 15 hours, and nothing crashed.
I also ran the Dell’s diagnostic tool, available in the boot menu of the motherboard. The whole test took over 2 hours but did not show any errors. It tested the memory, CPU, fans, storage drives, etc.

Sort:

Hot Top Controversial New Old

[ - ]

thesporkeffect@lemmy.world

4 points

5 months ago

I’d start here: https://askubuntu.com/questions/1264859/watchdog-bug-soft-lockup-cpu6-stuck-for-23s

permalink

report

[ - ]

bulwark@lemmy.world

7 points

5 months ago

I’ve never seen this particular error, but CPU stall warnings seem like a fairly common thing. I wouldn’t jump straight to hardware fault, but it’s a possibility.

https://docs.kernel.org/RCU/stallwarn.html

permalink

report

[ - ]

catloaf@lemm.ee

9 points

5 months ago

I’d lean toward bad hardware.

Try stress testing the CPU and RAM. See if you can get it to happen more frequently. Also see if you can disable that CPU core, either in the BIOS or in the OS, to see if the problem goes away.

permalink

report

parent

[ - ]

loganb@lemmy.world

8 points

5 months ago

I’m with catloaf. Consistent CPU soft locks point to a possible bad memory module or CPU.

Clear CMOS.

Try removing one memory module at a time.

See if there is an option to disable hyperthreading in bios.

Another thing to try is to remove the CPU, careful not to damage the LGA pins on the motherboard, and clean the CPU contacts with alcohol. Take care to ground yourself out and the case before handling the CPU out of socket.

permalink

report

parent

[ - ]

Possibly linux@lemmy.zip

3 points

5 months ago

Don’t try to clean CPU pins. That is a very bad idea

permalink

report

parent

Show more comments

[ - ]

just_another_person@lemmy.world

53 points

5 months ago

Seen this before, and almost always has to do with hardware failure or bad hardware config.

Reset the BIOS/CMOS jumper on the board, go back into BIOS setup and set the proper time. Do not touch the CPU or Memory timings. Boot with the defaults and see if it still happens. Check and update the BIOS if there is a newer version as well.

Next longer steps: test memory, then stress test the CPU. I’d be shocked if it was a storage issue as I haven’t seen that be the culprit, but might was well run the long SMART tests.

permalink

report

[ - ]

seaQueue@lemmy.world

20 points

5 months ago

Make sure the microcode package is up to date as well

permalink

report

parent

[ - ]

4am@lemm.ee

8 points

5 months ago

Yeah, always check all of this stuff. Server hardware gets a lot more updates than like gamer board BIOS, companies invest high millions, even low billions in this stuff and they expect problems to be address promptly for that kind of cash.

Check for any peripherals or cards, too. RAID, backplanes, networking cards; drivers, firmware, anything.

permalink

report

parent

[ - ]

Possibly linux@lemmy.zip

5 points

5 months ago

Poorly supported hardware will also do this

permalink

report

parent

[ - ]

seaQueue@lemmy.world

8 points

5 months ago

An Optiplex 5040 should be well and thoroughly supported for 6+y now

permalink

report

parent

[ - ]

nutbutter@discuss.tchncs.deOP

1 point

4 months ago

I did reset it. It did not help. I ran memtest86 for over 2 hours and did a CPU stress test for over 15 hours. Nothing crashed during the testing.

permalink

report

parent

[ - ]

ikidd@lemmy.world

9 points

5 months ago

Can you sudo dmesg | grep microcode and see if you have any errors?

If so, I’d be inclined to sudo dnf reinstall microcode_ctl then do a sudo dracut -f to regenerate the initramfs, and reboot. Be sure to have a working fallback kernel like LTS installed so you can recover if need be.

Edit: I just read you changed to Ubuntu. I can’t be arsed to figure out how Ubuntu does this stuff so that’s on you to figure out. Alternatively, install Arch and use intel-ucode package.

permalink

report

[ - ]

TheBigBrother@lemmy.world

-3 points

5 months ago

Set watchdog to reboot every day and juice it until the last drop before it definetly crashes.

Edit: that’s just a workaround if you really want try to fix it my answer would be restoring BIOS defaults, then clean install of the OS and then check it everything works fine, if the error persist try installing another OS if it still fail then go to the first step.

permalink

report

Selfhosted

!selfhosted@lemmy.world

Create post

A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don’t control.

Rules:

Be civil: we’re here to support and learn from one another. Insults won’t be tolerated. Flame wars are frowned upon.
No spam posting.
Posts have to be centered around self-hosting. There are other communities for discussing hardware or home computing. If it’s not obvious why your post topic revolves around selfhosting, please include details to make it clear.
Don’t duplicate the full text of your blog or github here. Just post the link for folks to click.
Submission headline should match the article title (don’t cherry-pick information from the title to fit your agenda).
No trolling.

Resources:

selfh.st Newsletter and index of selfhosted software and apps
awesome-selfhosted software
awesome-sysadmin resources
Self-Hosted Podcast from Jupiter Broadcasting

Any issues on the community? Report it using the report flag.

Questions? DM the mods!

Community stats

3.6K
Monthly active users
2K
Posts
23K
Comments

Community stats

Community moderators