Key Moments

High End Computing and Scientific Visualization at NASA

Google TalksGoogle Talks
Education6 min read63 min video
Aug 22, 2012|665 views|6
Save to Pod
TL;DR

NASA's Columbia supercomputer, at 62 teraflops, was the fastest operational system in 2004, enabling complex simulations for missions like Pluto and hurricane prediction, but requires vast storage and bandwidth to match its processing power.

Key Insights

1

NASA's Columbia supercomputer, with a peak performance of 62 teraflops, was ranked the second-fastest in the world in November 2004, running at 52 teraflops on the Linpack benchmark.

2

The Columbia system is composed of 20 SGI Altix 512-node machines, each a standalone 3-teraflop system, with a total of 10,240 processors and 1 terabyte of memory per node.

3

NASA has established a high-end computing revitalization task force and partnered with industry (Intel, SGI) to regain U.S. leadership in supercomputing, spurred by Japan's Earth Simulator and the Columbia accident.

4

The visualization group at NASA Ames supports research by developing reusable components like the 'field model' for data abstraction and the 'viztech' library for various visualization techniques.

5

The 'Growler' framework facilitates complex distributed visualization pipelines, managing components with different time scales and enabling interactive, real-time analysis of simulations.

6

A major challenge for Columbia is storage and bandwidth; the current 440 terabytes of storage is insufficient, and high-speed, agency-wide connectivity is needed to fully utilize the computational power.

The Columbia supercomputer: A leap in NASA's computational power

NASA's Columbia supercomputer, housed at the Ames Research Center, represented a significant advancement in high-end computing for the agency. Launched in 2004, it boasted a peak performance of 62 teraflops and achieved 52 teraflops on the Linpack benchmark, ranking it second globally at the time. The system is a massive aggregation of 20 SGI Altix 512-node machines, each a standalone single-system image shared-memory computer. Collectively, these provide approximately 10,240 processors, 1 terabyte of memory per node, and 20 terabytes of RAID storage per node. The impetus for developing such a powerful system stemmed from both external factors, like Japan's Earth Simulator, and internal needs, including the investigation of the Columbia accident and critical requirements for next-generation space flight modeling. This effort was part of a national initiative to restore U.S. leadership in high-end computing, involving collaboration between government agencies and industry partners like Intel and SGI.

Architectural details and key components of the Columbia system

The Columbia system's architecture is complex, built upon SGI Altix machines. It comprises various node types, including 1.5 GHz and 1.6 GHz processors with different cache sizes, and double-density configurations for tighter packing. These nodes are interconnected via multiple Gigabit interconnects, an InfiniBand fabric, and 10 Gigabit TCP interconnects. A crucial front-end system, a 128-processor Altix nicknamed 'Channel,' connects directly to graphics arrays, enabling real-time visualization. The system also includes substantial storage, totaling 440 terabytes of RAID disks, though this capacity is currently a bottleneck, not balancing the immense computational power. High-speed network connectivity, including efforts to leverage the National Lambda Rail, is essential for distributed data access across NASA's various centers.

Driving NASA's missions through high-performance computing

High-end computing at NASA is not merely about raw processing power; it's an integrated environment encompassing applications, storage, networks, and extensive support activities like post-processing and visualization. The Columbia system supports a diverse range of NASA's mission directorates, including aeronautics, exploration, science (earth and space), and space operations. Specific applications highlighted include detailed simulations for the New Horizons Pluto mission, return-to-flight efforts for the space shuttle main engine redesign, and cosmological studies for the Laser Interferometer Space Antenna (LISA) mission to understand the universe's origin by solving 3D Einstein equations for merging black holes. These simulations require significant computational effort and advanced visualization to interpret the results and guide future scientific inquiry.

Simulating extreme weather and climate phenomena

The Columbia supercomputer plays a vital role in advancing our understanding and prediction of Earth's climate and weather systems. For hurricane prediction, a key code runs on Columbia four times a day, feeding into ensemble models that combine predictions from various codes to forecast track, precipitation, and wind speeds. This work was crucial in improving hurricane forecasts, as demonstrated by its use in predicting Hurricane Ivan in 2004 and subsequent years. Furthermore, climate modeling efforts, such as a decadal run combining ocean and sea ice models using a cubic grid to avoid polar singularities, have revealed previously unseen features like vortices off the coast of South Africa. These complex simulations highlight the need for both powerful computation and sophisticated visualization techniques.

Visualization: Turning complex data into actionable insights

Chris Henze's group focuses on scientific visualization, supporting research by creating tools and techniques to interpret the vast datasets generated by NASA's supercomputers. Their work emphasizes interactive graphics and exploratory analysis, often visualizing time-varying phenomena and data from discretized meshes. Key challenges include handling terabyte-scale datasets and the need for 'concurrent visualization' to capture data as it's computed, minimizing input/output overhead. They have developed reusable components: the 'field model' (FM) for uniform data abstraction across different mesh types, the 'viztech' library for diverse visualization techniques (scalar, vector, feature detection), and the 'Growler' distributed framework for managing complex visualization pipelines across disparate compute and display systems.

The 'Field Model' and 'Viztech' for data analysis

The 'field model' (FM) provides a uniform interface to underlying discrete data, treating computational domains as continuous fields. It allows users to query data values at arbitrary locations, performing interpolation and derivation where needed. Optimizations include exploiting symmetries by using reference meshes and employing paging to load only necessary data chunks. The 'viztech' library encapsulates various visualization techniques, ranging from visualizing mesh geometry to scalar and vector fields, and advanced feature detection algorithms like vortex core finders. These tools are essential for scientists to probe data, identify critical features, and gain a deeper understanding of complex phenomena, such as flow reversal in engine fuel injection systems or the evolution of large-scale structures in the early universe.

The 'Growler' framework and interactive visualization environments

The 'Growler' framework is designed to connect disparate components—data sources, compute engines, and display devices—that often operate on different time scales. It uses a remote method invocation (RMI) service, an interface definition language, and a sophisticated signal selector mechanism to manage complex distributed pipelines. This enables interactive and immersive visualization experiences, such as computational steering of molecular dynamics simulations with haptic feedback. For hurricane forecasting, 'Growler' facilitated real-time interface with live simulations, processing data and streaming visualizations for analysis. The group also utilizes display systems like the 'Hyperwall,' a large array of flat panel displays, for visualizing large datasets and exploring high-dimensional dependencies through linked views and parameterized layouts.

Future challenges: Storage, bandwidth, and metadata

Despite the advancements, significant challenges remain. The computational power of Columbia is not yet matched by its storage capacity; 440 terabytes is insufficient for the volume of data generated, necessitating much more storage. Similarly, achieving end-to-end distributed data access across NASA centers requires high-speed, high-bandwidth connectivity, a project still in progress. Furthermore, managing the metadata associated with large binary objects and enabling efficient searching over this data is an ongoing interest. The complexity of these distributed systems also presents debugging challenges, often requiring a modular, divide-and-conquer approach, with failures typically occurring in large distributed components or integration issues.

High-End Computing & Visualization Best Practices at NASA

Practical takeaways from this episode

Do This

Leverage national partnerships and agency cooperation for diverse computing architectures.
Collaborate closely with scientists to understand requirements and develop solutions.
Integrate storage and high-speed networks with computational power for efficient data handling.
Utilize visualization tools to analyze large datasets and identify key features.
Employ concurrent visualization to process data as it's being computed, minimizing I/O.
Exploit symmetries and paging optimizations in data models for memory efficiency.
Use distributed frameworks to manage components with different time scales.
Consider hybrid MPI+OpenMP or MLP programming models for parallelization.
Leverage tile display systems like Hyperwall for comparing parameterized layouts and linked visualizations.

Avoid This

Do not assume one architecture fits all applications.
Do not overlook the importance of integrated environments including storage, networks, and support activities.
Do not neglect the need for high-bandwidth connectivity across distributed agency centers.
Do not rely solely on raw computational power; balance it with storage and network capabilities.
Do not wait until simulation completion to visualize; use concurrent visualization.
Do not ignore the challenges of debugging complex distributed systems.
Do not limit visualization to static images; use animations and interactive analysis.

Common Questions

High-end computing at NASA refers to the use of powerful supercomputers like the Columbia system for complex simulations and data analysis across various mission directorates, including aeronautics, exploration, earth science, and space operations.

Topics

Mentioned in this video

Organizations
NASA

The U.S. federal agency responsible for the civilian space program, as well as aeronautics and space research. The video discusses its organization, mission directorates, and high-end computing efforts.

JPL

Jet Propulsion Laboratory, mentioned as having connections to Columbia.

Pew Research Center

One of NASA's research centers, specifically mentioned as the location of the Columbia machine and where the presentation is being given.

Johnson Space Center

A NASA center with specific missions.

National Lambda Rail

A network used for high-speed connectivity to Columbia, a consortium of industry, academic partners, and other agencies.

Kennedy Space Center

A NASA center with specific missions.

University of Florida

Where data from the hurricane prediction code run on Columbia is sent for further processing with other ensemble models.

Max Planck Institute

A message-passing interface standard for distributed-memory parallelism, discussed as a programming model for the Columbia machine, often used in hybrid mode with OpenMP.

Langley Research Center

One of NASA's research centers.

Goddard Space Flight Center

A NASA center focused on Earth sciences.

National Coordination Office

An office within the US government that established the High-End Computing Revitalization Task Force.

High-End Computing Revitalization Task Force

A task force formed by the US government to restore US leadership in high-end computing.

LISA

A future mission (planned for 2015) between NASA and ESA to study the origin of the universe by measuring gravitational waves from merging black holes.

European Space Agency

Collaborated with NASA on the LISA mission.

National Hurricane Center

The U.S. agency responsible for hurricane predictions, which uses data from models like those run on Columbia.

More from GoogleTalksArchive

View all 83 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free