High Performance Parallelism Pearls Volume One: Multicore and Many-core Programming Approaches

High Performance Parallelism Pearls Volume One: Multicore and Many-core Programming Approaches

High Performance Parallelism Pearls Volume One: Multicore and Many-core Programming Approaches

High Performance Parallelism Pearls Volume One: Multicore and Many-core Programming Approaches

eBook

$59.99  $79.95 Save 25% Current price is $59.99, Original price is $79.95. You Save 25%.

Available on Compatible NOOK devices, the free NOOK App and in My Digital Library.
WANT A NOOK?  Explore Now

Related collections and offers


Overview

High Performance Parallelism Pearls shows how to leverage parallelism on processors and coprocessors with the same programming – illustrating the most effective ways to better tap the computational potential of systems with Intel Xeon Phi coprocessors and Intel Xeon processors or other multicore processors. The book includes examples of successful programming efforts, drawn from across industries and domains such as chemistry, engineering, and environmental science. Each chapter in this edited work includes detailed explanations of the programming techniques used, while showing high performance results on both Intel Xeon Phi coprocessors and multicore processors. Learn from dozens of new examples and case studies illustrating "success stories" demonstrating not just the features of these powerful systems, but also how to leverage parallelism across these heterogeneous systems.

  • Promotes consistent standards-based programming, showing in detail how to code for high performance on multicore processors and Intel® Xeon Phi™
  • Examples from multiple vertical domains illustrating parallel optimizations to modernize real-world codes
  • Source code available for download to facilitate further exploration

Product Details

ISBN-13: 9780128021996
Publisher: Elsevier Science
Publication date: 11/04/2014
Sold by: Barnes & Noble
Format: eBook
Pages: 600
File size: 61 MB
Note: This product may take a few minutes to download.

About the Author

James Reinders is a senior engineer who joined Intel Corporation in 1989 and has contributed to projects including the world’s first TeraFLOP supercomputer (ASCI Red), as well as compilers and architecture work for a number of Intel processors and parallel systems. James has been a driver behind the development of Intel as a major provider of software development products, and serves as their chief software evangelist. James has published numerous articles, contributed to several books and is widely interviewed on parallelism. James has managed software development groups, customer service and consulting teams, business development and marketing teams. James is sought after to keynote on parallel programming, and is the author/co-author of three books currently in print including Structured Parallel Programming, published by Morgan Kaufmann in 2012.
Jim Jeffers was the primary strategic planner and one of the first full-time employees on the program that became Intel ® MIC. He served as lead SW Engineering Manager on the program and formed and launched the SW development team. As the program evolved, he became the workloads (applications) and SW performance team manager. He has some of the deepest insight into the market, architecture and programming usages of the MIC product line. He has been a developer and development manager for embedded and high performance systems for close to 30 years.

Table of Contents

1. Introduction2. Towards an efficient Godunov's scheme on Phi3. Better Concurrency and SIMD on HBM4. Case Study: Analyzing and Optimizing Concurrency5. Plesiochronous Phasing Barriers6. Parallel Evaluation of Fault Tree Expressions7. Deep-learning and Numerical Optimization8. Optimizing Gather/Scatter Patterns9. A many core implementation of the direct N-body problem10. N-body Methods on Intel® Xeon Phi™ Coprocessors11. Dynamic Load Balancing using OpenMP 4.012. Concurrent Kernel Offloading13. Heterogeneous Computing with MPI14. Power Analysis on the Intel® Xeon Phi™ Coprocessor15. Integrating Intel Xeon Phis into a Cluster16. Native File systems17. NWChem: Quantum Chemistry Simulations at Scale18. Efficient nested parallelism on large scale system19. Performance optimization of Black-Scholes pricing20. Host and Coprocessor Data Transfer through the COI21. High Performance Ray Tracing with Embree22. Portable and Perform with OpenCL23. Characterization and Auto-tuning of 3DFD.24. Profiling-guided optimization of cache performance25. Heterogeneous MPI optimization with ITAC26. Scalable Out-of-core Solvers on a Cluster27. Sparse matrix-vector multiplication: parallelization and vectorization28. Morton Order Improves Performance

What People are Saying About This

From the Publisher

Case studies and examples illustrating the power of high performance parallelism

From the B&N Reads Blog

Customer Reviews