


{"id":382381,"date":"2024-09-11T17:08:05","date_gmt":"2024-09-11T17:08:05","guid":{"rendered":"https:\/\/innovationspace.ansys.com\/forum\/forums\/reply\/382381\/"},"modified":"2024-09-11T17:08:05","modified_gmt":"2024-09-11T17:08:05","slug":"382381","status":"publish","type":"reply","link":"https:\/\/innovationspace.ansys.com\/forum\/forums\/reply\/382381\/","title":{"rendered":"Reply To: Computation Accleration"},"content":{"rendered":"<p>&lt;p&gt;Hi Zihan&lt;\/p&gt;&lt;p&gt;For the first question let me give a little backgorund first.&nbsp; In distributed parallel solutions, the FEM domain is split into N domains at the element level if using N cpu cores.&nbsp; Then N instances of the solver process (here MAPDL) are used to solve and each solve process takes one of the smaller domains to solve.&nbsp; Each process has to communicate with the other processes that share nodes (FEM nodes, not compute cluster nodes).&nbsp; If the amount of communication gets to be more than the computation that the cpu core is doing, then too many CPU cores are being used.&nbsp; So for any model there will be a point where using more CPU cores will just slow down the solve.&lt;\/p&gt;&lt;p&gt;In the mathematics of solving FEA equations we use degrees of freedom [DOF] instead of number of nodes\/elements.&nbsp; Let&#8217;s say this is a FEM with only solid structural elements with translational DOFs in X, Y and Z.&nbsp; Then the total number of DOFs is 3 * number of nodes.&lt;\/p&gt;&lt;p&gt;In order to know how many cpu cores to use we need to know the range of DOFs per cpu core where the cpu core is most efficient.&nbsp; When working with a new CPU that I&#8217;ve not tested before I assume something like 30,000 DOFs per cpu core and figure out how many cpus and compute nodes are needed from there.&nbsp; Ignoring any question on network communication speed between the compute nodes (for now).&lt;\/p&gt;&lt;p&gt;So for your model I&#8217;d try 170 cpu cores or 7 compute nodes as a test.&nbsp; Then maybe try 6 and 8 compute nodes (fully used) to see how hardware responds to a little more and less dof&#8217;s per cpu core.&nbsp; Depending on what happens you may then want to run other tests.&lt;\/p&gt;&lt;p&gt;For now I&#8217;d not use the GPU as a solver accelerator.&nbsp; Since the GPU is helping all of the processes one GPU for 24 cpu cores is a bit much&#8230;I&#8217;d prefer to see 2 gpus for those 24 cores.&nbsp; Trying to keep to about 1 gpu per 12 or so cpu cores (and hence number of solver processes).&lt;\/p&gt;&lt;p&gt;I&#8217;m also assuming the use of the sparse (direct) solver.&nbsp; If using the iterative solver then the dof&#8217;s per cpu core will probably be different for that cluster.&lt;\/p&gt;&lt;p&gt;&nbsp;&lt;\/p&gt;<\/p>\n","protected":false},"template":"","class_list":["post-382381","reply","type-reply","status-publish","hentry"],"aioseo_notices":[],"aioseo_head":"\n\t\t<!-- All in One SEO 4.9.10 - aioseo.com -->\n\t<meta name=\"description\" content=\"Hi ZihanFor the first question let me give a little backgorund first. In distributed parallel solutions, the FEM domain is split into N domains at the element level if using N cpu cores. Then N instances of the solver process (here MAPDL) are used to solve and each solve process takes one of the smaller\" \/>\n\t<meta name=\"robots\" content=\"max-image-preview:large\" \/>\n\t<link rel=\"canonical\" href=\"https:\/\/innovationspace.ansys.com\/forum\/forums\/reply\/382381\/\" \/>\n\t<meta name=\"generator\" content=\"All in One SEO (AIOSEO) 4.9.10\" \/>\n\t\t<meta property=\"og:locale\" content=\"en_US\" \/>\n\t\t<meta property=\"og:site_name\" content=\"Ansys Learning Forum | Ansys Innovation Space\" \/>\n\t\t<meta property=\"og:type\" content=\"article\" \/>\n\t\t<meta property=\"og:title\" content=\"Reply To: Computation Accleration | Ansys Learning Forum\" \/>\n\t\t<meta property=\"og:description\" content=\"Hi ZihanFor the first question let me give a little backgorund first. In distributed parallel solutions, the FEM domain is split into N domains at the element level if using N cpu cores. Then N instances of the solver process (here MAPDL) are used to solve and each solve process takes one of the smaller\" \/>\n\t\t<meta property=\"og:url\" content=\"https:\/\/innovationspace.ansys.com\/forum\/forums\/reply\/382381\/\" \/>\n\t\t<meta property=\"article:published_time\" content=\"2024-09-11T17:08:05+00:00\" \/>\n\t\t<meta property=\"article:modified_time\" content=\"2024-09-11T17:08:05+00:00\" \/>\n\t\t<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n\t\t<meta name=\"twitter:title\" content=\"Reply To: Computation Accleration | Ansys Learning Forum\" \/>\n\t\t<meta name=\"twitter:description\" content=\"Hi ZihanFor the first question let me give a little backgorund first. In distributed parallel solutions, the FEM domain is split into N domains at the element level if using N cpu cores. Then N instances of the solver process (here MAPDL) are used to solve and each solve process takes one of the smaller\" \/>\n\t\t<script type=\"application\/ld+json\" class=\"aioseo-schema\">\n\t\t\t{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/forums\\\/reply\\\/382381\\\/#breadcrumblist\",\"itemListElement\":[{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum#listItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\",\"nextItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/forums\\\/reply\\\/382381\\\/#listItem\",\"name\":\"Reply To: Computation Accleration\"}},{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/forums\\\/reply\\\/382381\\\/#listItem\",\"position\":2,\"name\":\"Reply To: Computation Accleration\",\"previousItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum#listItem\",\"name\":\"Home\"}}]},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/#organization\",\"name\":\"Ansys Learning Forum\",\"description\":\"Ansys Innovation Space\",\"url\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/forums\\\/reply\\\/382381\\\/#webpage\",\"url\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/forums\\\/reply\\\/382381\\\/\",\"name\":\"Reply To: Computation Accleration | Ansys Learning Forum\",\"description\":\"Hi ZihanFor the first question let me give a little backgorund first. In distributed parallel solutions, the FEM domain is split into N domains at the element level if using N cpu cores. Then N instances of the solver process (here MAPDL) are used to solve and each solve process takes one of the smaller\",\"inLanguage\":\"en-US\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/#website\"},\"breadcrumb\":{\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/forums\\\/reply\\\/382381\\\/#breadcrumblist\"},\"datePublished\":\"2024-09-11T17:08:05+00:00\",\"dateModified\":\"2024-09-11T17:08:05+00:00\"},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/#website\",\"url\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/\",\"name\":\"Ansys Learning Forum\",\"description\":\"Ansys Innovation Space\",\"inLanguage\":\"en-US\",\"publisher\":{\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/#organization\"}}]}\n\t\t<\/script>\n\t\t<!-- All in One SEO -->\n\n","aioseo_head_json":{"title":"Reply To: Computation Accleration | Ansys Learning Forum","description":"Hi ZihanFor the first question let me give a little backgorund first. In distributed parallel solutions, the FEM domain is split into N domains at the element level if using N cpu cores. Then N instances of the solver process (here MAPDL) are used to solve and each solve process takes one of the smaller","canonical_url":"https:\/\/innovationspace.ansys.com\/forum\/forums\/reply\/382381\/","robots":"max-image-preview:large","keywords":"","webmasterTools":{"miscellaneous":""},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"BreadcrumbList","@id":"https:\/\/innovationspace.ansys.com\/forum\/forums\/reply\/382381\/#breadcrumblist","itemListElement":[{"@type":"ListItem","@id":"https:\/\/innovationspace.ansys.com\/forum#listItem","position":1,"name":"Home","item":"https:\/\/innovationspace.ansys.com\/forum","nextItem":{"@type":"ListItem","@id":"https:\/\/innovationspace.ansys.com\/forum\/forums\/reply\/382381\/#listItem","name":"Reply To: Computation Accleration"}},{"@type":"ListItem","@id":"https:\/\/innovationspace.ansys.com\/forum\/forums\/reply\/382381\/#listItem","position":2,"name":"Reply To: Computation Accleration","previousItem":{"@type":"ListItem","@id":"https:\/\/innovationspace.ansys.com\/forum#listItem","name":"Home"}}]},{"@type":"Organization","@id":"https:\/\/innovationspace.ansys.com\/forum\/#organization","name":"Ansys Learning Forum","description":"Ansys Innovation Space","url":"https:\/\/innovationspace.ansys.com\/forum\/"},{"@type":"WebPage","@id":"https:\/\/innovationspace.ansys.com\/forum\/forums\/reply\/382381\/#webpage","url":"https:\/\/innovationspace.ansys.com\/forum\/forums\/reply\/382381\/","name":"Reply To: Computation Accleration | Ansys Learning Forum","description":"Hi ZihanFor the first question let me give a little backgorund first. In distributed parallel solutions, the FEM domain is split into N domains at the element level if using N cpu cores. Then N instances of the solver process (here MAPDL) are used to solve and each solve process takes one of the smaller","inLanguage":"en-US","isPartOf":{"@id":"https:\/\/innovationspace.ansys.com\/forum\/#website"},"breadcrumb":{"@id":"https:\/\/innovationspace.ansys.com\/forum\/forums\/reply\/382381\/#breadcrumblist"},"datePublished":"2024-09-11T17:08:05+00:00","dateModified":"2024-09-11T17:08:05+00:00"},{"@type":"WebSite","@id":"https:\/\/innovationspace.ansys.com\/forum\/#website","url":"https:\/\/innovationspace.ansys.com\/forum\/","name":"Ansys Learning Forum","description":"Ansys Innovation Space","inLanguage":"en-US","publisher":{"@id":"https:\/\/innovationspace.ansys.com\/forum\/#organization"}}]},"og:locale":"en_US","og:site_name":"Ansys Learning Forum | Ansys Innovation Space","og:type":"article","og:title":"Reply To: Computation Accleration | Ansys Learning Forum","og:description":"Hi ZihanFor the first question let me give a little backgorund first. In distributed parallel solutions, the FEM domain is split into N domains at the element level if using N cpu cores. Then N instances of the solver process (here MAPDL) are used to solve and each solve process takes one of the smaller","og:url":"https:\/\/innovationspace.ansys.com\/forum\/forums\/reply\/382381\/","article:published_time":"2024-09-11T17:08:05+00:00","article:modified_time":"2024-09-11T17:08:05+00:00","twitter:card":"summary_large_image","twitter:title":"Reply To: Computation Accleration | Ansys Learning Forum","twitter:description":"Hi ZihanFor the first question let me give a little backgorund first. In distributed parallel solutions, the FEM domain is split into N domains at the element level if using N cpu cores. Then N instances of the solver process (here MAPDL) are used to solve and each solve process takes one of the smaller"},"aioseo_meta_data":{"post_id":"382381","title":null,"description":null,"keywords":null,"keyphrases":null,"primary_term":null,"canonical_url":null,"og_title":null,"og_description":null,"og_object_type":"default","og_image_type":"content","og_image_url":null,"og_image_width":null,"og_image_height":null,"og_image_custom_url":null,"og_image_custom_fields":null,"og_video":null,"og_custom_url":null,"og_article_section":null,"og_article_tags":null,"twitter_use_og":true,"twitter_card":"default","twitter_image_type":"default","twitter_image_url":null,"twitter_image_custom_url":null,"twitter_image_custom_fields":null,"twitter_title":null,"twitter_description":null,"schema":{"blockGraphs":[],"customGraphs":[],"default":{"data":{"Article":[],"Course":[],"Dataset":[],"FAQPage":[],"Movie":[],"Person":[],"Product":[],"ProductReview":[],"Car":[],"Recipe":[],"Service":[],"SoftwareApplication":[],"WebPage":[]},"graphName":"","isEnabled":true},"graphs":[]},"schema_type":"default","schema_type_options":null,"pillar_content":false,"robots_default":true,"robots_noindex":false,"robots_noarchive":false,"robots_nosnippet":false,"robots_nofollow":false,"robots_noimageindex":false,"robots_noodp":false,"robots_notranslate":false,"robots_max_snippet":null,"robots_max_videopreview":null,"robots_max_imagepreview":"large","priority":null,"frequency":null,"local_seo":null,"limit_modified_date":false,"created":"2024-10-28 16:29:24","updated":"2026-08-16 03:20:14","ai":null,"breadcrumb_settings":null,"seo_analyzer_scan_date":null},"aioseo_breadcrumb":"<div class=\"aioseo-breadcrumbs\"><span class=\"aioseo-breadcrumb\">\n\t\t\t<a href=\"https:\/\/innovationspace.ansys.com\/forum\" title=\"Home\">Home<\/a>\n\t\t<\/span><span class=\"aioseo-breadcrumb-separator\">&raquo;<\/span><span class=\"aioseo-breadcrumb\">\n\t\t\tReply To: Computation Accleration\n\t\t<\/span><\/div>","aioseo_breadcrumb_json":[{"label":"Home","link":"https:\/\/innovationspace.ansys.com\/forum"},{"label":"Reply To: Computation Accleration","link":"https:\/\/innovationspace.ansys.com\/forum\/forums\/reply\/382381\/"}],"acf":[],"_links":{"self":[{"href":"https:\/\/innovationspace.ansys.com\/forum\/wp-json\/wp\/v2\/replies\/382381","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/innovationspace.ansys.com\/forum\/wp-json\/wp\/v2\/replies"}],"about":[{"href":"https:\/\/innovationspace.ansys.com\/forum\/wp-json\/wp\/v2\/types\/reply"}],"version-history":[{"count":0,"href":"https:\/\/innovationspace.ansys.com\/forum\/wp-json\/wp\/v2\/replies\/382381\/revisions"}],"wp:attachment":[{"href":"https:\/\/innovationspace.ansys.com\/forum\/wp-json\/wp\/v2\/media?parent=382381"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}