CAVU Aerospace UK

SEU Resilience in the OBC-Hyper-Polar

A Layered Fault-Tolerant Architecture for Reliable Space Computing

Spacecraft electronics operate in one of the harshest environments ever encountered by digital systems. Unlike terrestrial computers, onboard computers (OBCs) are continuously exposed to energetic protons, heavy ions, cosmic rays and trapped particles capable of disrupting electronic circuits. While modern semiconductor technology has significantly improved performance and reduced power consumption, it has also increased susceptibility to radiation-induced faults.

Among these radiation effects, Single Event Upsets are the most common and one of the primary design challenges for spacecraft avionics.

The OBC-Hyper-Polar has been engineered with the assumption that radiation events are inevitable. Instead of relying solely on radiation-hardened components, the platform employs a comprehensive Fault Detection, Isolation and Recovery (FDIR) architecture that prevents, detects, isolates and autonomously recovers from radiation-induced faults. This layered approach enables reliable operation across a wide range of missions, from commercial CubeSats to institutional and scientific spacecraft.

A Single Event Upset is a temporary change in the state of a digital circuit caused by a single energetic particle passing through semiconductor material. As the particle traverses the silicon, it deposits charge that may alter the logic state of a memory cell, processor register or flip-flop.

Unlike permanent radiation damage, an SEU does not physically damage the device. It is therefore known as a soft error. Nevertheless, an undetected SEU can lead to incorrect computations, corrupted data or unexpected system behaviour.

Typical consequences are corrupted memory contents, processor exceptions, communication errors, software crashes & unexpected subsystem resets.

The probability of SEUs depends on orbital altitude, inclination, mission duration, solar activity and spacecraft shielding. Even satellites operating in Low Earth Orbit experience regular radiation events, particularly during passages through the South Atlantic Anomaly and periods of elevated solar activity.

Although SEUs receive the greatest attention, spacecraft designers must also consider other radiation phenomena.

  • Single Event Transient (SET): A temporary voltage pulse that may propagate through combinational logic.
  • Single Event Functional Interrupt (SEFI): A radiation-induced interruption requiring subsystem reinitialisation.
  • Single Event Latch-up (SEL): A potentially destructive condition causing excessive current consumption.
  • Multiple Bit Upsets (MBU): Simultaneous corruption of multiple memory bits.
  • Total Ionizing Dose (TID): Long-term degradation caused by cumulative radiation exposure.

Each phenomenon requires different mitigation techniques. Consequently, reliable spacecraft design cannot depend on a single technology or component.

 

Flash-Based FPGA: Eliminating Configuration SEUs

At the heart of every OBC-Hyper-Polar is a Microchip PolarFire FPGA SoC. Unlike conventional SRAM-based FPGAs, these devices store their configuration in non-volatile Flash memory. This architecture provides a major reliability advantage because the FPGA configuration itself is inherently resistant to radiation-induced configuration upsets.

As a result, configuration scrubbing is not required & the implemented hardware design cannot be altered by a transient particle strike. Power consumption associated with continuous scrubbing is eliminated & system complexity is significantly reduced. This allows the fault-tolerant architecture to focus on protecting volatile resources such as processor registers, SRAM, DDR memory and communication interfaces.

SEU Resilience, OBC-Hyper-Polar, Layered Fault-Tolerant Architecture, Reliable Space Computing, Spacecraft electronics, terrestrial computers, onboard computers, OBC, radiation-induced faults, Single Event Upsets, spacecraft avionics, radiation-hardened components, FDIR, commercial CubeSats, spacecraft, permanent radiation damage, orbital altitude, spacecraft shielding, Low Earth Orbit, Single Event Transient, Single Event Functional Interrupt, Single Event Latch-up, Multiple Bit Upsets, Total Ionizing Dose, Flash-Based FPGA, Microchip PolarFire FPGA SoC, SRAM-based FPGAs, Layered SEU Protection Strategy, Fault Detection, Isolation and Recovery, COTS version, LEO missions, RT version, Radiation-Tolerant

A Layered SEU Protection Strategy

The OBC-Hyper-Polar does not depend on any single mitigation technique. Instead, multiple independent protection layers operate simultaneously throughout the system.

The first layer reduces the likelihood of radiation-induced failures through careful component selection and robust hardware architecture. Depending on the mission profile, the platform incorporates flash-based FPGA, radiation-tolerant components for critical subsystems, redundant power converters, protected power distribution & mission-specific component selection.

Detection- Since transient faults cannot be completely eliminated, every critical subsystem is continuously monitored. The onboard computer implements:

  • Hardware watchdog timers
  • Software watchdog supervision
  • ECC error detection
  • CRC verification
  • SHA-256 boot image validation
  • TMR voter mismatch monitoring
  • Voltage and current monitoring
  • Temperature monitoring
  • Communication error counters
  • Housekeeping telemetry
  • Continuous health monitoring

These mechanisms enable rapid identification of abnormal system behaviour before faults propagate throughout the spacecraft.

Isolation- Once a fault has been detected, the affected subsystem must be isolated to prevent cascading failures. The FDIR architecture supports:

  • Faulty communication channel isolation
  • Power rail isolation using electronic fuses
  • Memory bank isolation
  • Redundant storage selection
  • Processor subsystem restart
  • Safe-mode transition for critical failures

This approach ensures that local faults do not compromise the entire spacecraft.

Recovery- Autonomous recovery is one of the defining characteristics of the OBC-Hyper-Polar architecture. Recovery mechanisms include:

  • Automatic processor restart
  • FPGA subsystem re-initialisation
  • Memory reconstruction using ECC
  • Recovery from redundant boot images
  • Automatic communication interface recovery
  • Switching to redundant storage devices
  • Controlled transition to spacecraft safe mode

Mission operations therefore continue with minimal interruption while reducing dependence on immediate ground intervention.

 

FDIR: The Core of System Resilience

The Fault Detection, Isolation and Recovery architecture provides continuous supervision of every critical subsystem. The design philosophy assumes that transient faults will occur during every sufficiently long mission. Rather than attempting to prevent every upset, the FDIR continuously evaluates system health and executes predefined recovery actions whenever abnormal behaviour is detected. Examples include:

  • Hardware watchdog reset of the processor subsystem.
  • TMR voter mismatch detection within FPGA logic.
  • ECC correction of single-bit memory errors.
  • Periodic memory integrity verification.
  • Automatic boot image validation using SHA-256.
  • Redundant boot image selection following corruption.
  • CRC verification of non-volatile memory.
  • Automatic failover between redundant storage devices.
  • Recovery of communication interfaces after persistent transmission errors.

Each fault therefore has a defined detection mechanism, isolation strategy and recovery sequence.

 

SEU Mitigation Throughout the OBC

Protection is integrated into every major subsystem rather than concentrated in a single device.

Component

SEU Mitigation

FPGA Configuration

Flash-based configuration memory inherently resistant to configuration SEUs

FPGA Logic

Triple Modular Redundancy (where implemented), voter monitoring and autonomous recovery

Processor

Hardware watchdog, software supervision and autonomous restart

DDR Memory

SECDED ECC, periodic integrity verification and memory recovery

Boot Memory

Dual QSPI images with SHA-256 integrity verification

MRAM

CRC protection with redundant storage architecture

eMMC Storage

Dual-device redundancy with automatic failover

Power System

Dual converters, electronic fuses and continuous monitoring

Communications

CRC, error counters, watchdog monitoring and interface recovery

This distributed approach significantly increases overall system availability by ensuring that failures remain localised and recoverable.

Component

Detection

Recovery

Power input

eFuse fault flags

Automatic isolation and retry

5 V power

Voltage monitoring

Redundant converter takeover

PolarFire MSS

Hardware watchdog

Autonomous processor reset

FPGA logic

TMR mismatch counters

Self-correction and reinitialization

DDR4

SECDED ECC

Memory correction or restart

eMMC

ECC and CRC

Automatic switch to redundant device

QSPI boot

SHA-256 verification

Boot from redundant image

MRAM

CRC

Alternate bank recovery

Ethernet

Link monitoring

Automatic port failover

CAN/RS-422/RS-485

Error counters

Redundant channel selection

 

OBC-Hyper-Polar COTS version

The COTS version is intended for commercial spacecraft, rockets, technology demonstrators and short-to-medium duration LEO missions. Although based on commercial components, the architecture incorporates the same layered FDIR philosophy used throughout the product family, including:

  • Flash-based FPGA configuration immunity
  • ECC-protected memories
  • Hardware and software watchdogs
  • CRC and integrity monitoring
  • Redundant boot architecture
  • Autonomous recovery mechanisms

This provides an excellent balance between reliability, performance and cost.

 

OBC-Hyper-Polar RT version

The Radiation-Tolerant version extends reliability through selective replacement of mission-critical components with radiation-tolerant alternatives. Rather than replacing every component, engineering analysis identifies those subsystems that contribute most significantly to mission risk. Critical memories, power management devices and interface components are upgraded where they provide measurable improvements in spacecraft reliability. The result is a substantial increase in radiation tolerance while maintaining competitive cost and avoiding unnecessary complexity.

For long-duration institutional, scientific and exploration missions, the RT configuration provides the highest level of radiation assurance. This version may incorporate:

  • Radiation-tolerant PolarFire FPGA devices
  • Radiation-tolerant memories
  • Enhanced redundancy
  • Increased Total Ionizing Dose capability
  • Mission-specific radiation analysis
  • Customised FDIR configuration
  • Extended environmental qualification

The RT version is intended for missions where maximum availability and operational lifetime are primary design objectives with any given costs. So we have several tier of radiation tolerance in OBC-Hyper-Polar which gives different levels of reliability in different costs. We have identified critical components to have proper reliability cost/ worth analysis.