Assign priorities and responsible teams to each node. Then map dependencies so one failure doesn't trigger hundreds of alerts.
Only then, re-enable basic up/down alerts with notification methods matched to each priority level (e.g. P1 sends an SMS, P4 sends an email, P5 appears on a dashboard).
This way, every alert that fires is trustworthy.

