


{"id":362344,"date":"2024-04-15T22:44:24","date_gmt":"2024-04-15T22:44:24","guid":{"rendered":"\/forum\/forums\/topic\/lumerical-cluster-job-crashes-arbitrarily\/"},"modified":"2024-04-15T22:44:24","modified_gmt":"2024-04-15T22:44:24","slug":"lumerical-cluster-job-crashes-arbitrarily","status":"closed","type":"topic","link":"https:\/\/innovationspace.ansys.com\/forum\/forums\/topic\/lumerical-cluster-job-crashes-arbitrarily\/","title":{"rendered":"Lumerical cluster job crashes arbitrarily"},"content":{"rendered":"<p>I run many parallel jobs on a cluster. Sometimes &#8211; for no apparent reason, the job crashes citing &#8220;std::bad_alloc&#8221;, but no more information. The RAM is sufficient (I have verified this, I have allocated 10x the RAM that the requirements ask for &#8211; and have monitored that this is not reached via observing using htop). This crash doesnt happen if I use just a single core on a local computer resource setting. However it happens on parallel slurm jobs with multiple cores. I need to use multiple cores to speed up simulations.<\/p>\n<p>This crash is not reproducible &#8211; it randomly happens, and if I run the same job again (same simulation with same resources in the resource manager), it sometimes runs and sometimes doesnt. Thus this leads me to believe that it doesn&#8217;t have anything to do with RAM but is some other problem.<\/p>\n<p>This issue is okay sometimes &#8211; where for many job submissions it does not crash, but sometimes it crashes quite often, which is not desirable &#8211; and it interrupts sweeps of simulations that I run.&nbsp;<\/p>\n<p>This is with Lumerical version 2022 R2<\/p>\n","protected":false},"template":"","class_list":["post-362344","topic","type-topic","status-closed","hentry"],"aioseo_notices":[],"aioseo_head":"\n\t\t<!-- All in One SEO 4.9.10 - aioseo.com -->\n\t<meta name=\"description\" content=\"I run many parallel jobs on a cluster. Sometimes - for no apparent reason, the job crashes citing &quot;std::bad_alloc&quot;, but no more information. The RAM is sufficient (I have verified this, I have allocated 10x the RAM that the requirements ask for - and have monitored that this is not reached via observing using htop).\" \/>\n\t<meta name=\"robots\" content=\"max-image-preview:large\" \/>\n\t<link rel=\"canonical\" href=\"https:\/\/innovationspace.ansys.com\/forum\/forums\/topic\/lumerical-cluster-job-crashes-arbitrarily\/\" \/>\n\t<meta name=\"generator\" content=\"All in One SEO (AIOSEO) 4.9.10\" \/>\n\t\t<meta property=\"og:locale\" content=\"en_US\" \/>\n\t\t<meta property=\"og:site_name\" content=\"Ansys Learning Forum | Ansys Innovation Space\" \/>\n\t\t<meta property=\"og:type\" content=\"article\" \/>\n\t\t<meta property=\"og:title\" content=\"Lumerical cluster job crashes arbitrarily | Ansys Learning Forum\" \/>\n\t\t<meta property=\"og:description\" content=\"I run many parallel jobs on a cluster. Sometimes - for no apparent reason, the job crashes citing &quot;std::bad_alloc&quot;, but no more information. The RAM is sufficient (I have verified this, I have allocated 10x the RAM that the requirements ask for - and have monitored that this is not reached via observing using htop).\" \/>\n\t\t<meta property=\"og:url\" content=\"https:\/\/innovationspace.ansys.com\/forum\/forums\/topic\/lumerical-cluster-job-crashes-arbitrarily\/\" \/>\n\t\t<meta property=\"article:published_time\" content=\"2024-04-15T22:44:24+00:00\" \/>\n\t\t<meta property=\"article:modified_time\" content=\"2024-04-15T22:44:24+00:00\" \/>\n\t\t<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n\t\t<meta name=\"twitter:title\" content=\"Lumerical cluster job crashes arbitrarily | Ansys Learning Forum\" \/>\n\t\t<meta name=\"twitter:description\" content=\"I run many parallel jobs on a cluster. Sometimes - for no apparent reason, the job crashes citing &quot;std::bad_alloc&quot;, but no more information. The RAM is sufficient (I have verified this, I have allocated 10x the RAM that the requirements ask for - and have monitored that this is not reached via observing using htop).\" \/>\n\t\t<script type=\"application\/ld+json\" class=\"aioseo-schema\">\n\t\t\t{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/forums\\\/topic\\\/lumerical-cluster-job-crashes-arbitrarily\\\/#breadcrumblist\",\"itemListElement\":[{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum#listItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\",\"nextItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/topics\\\/#listItem\",\"name\":\"Topics\"}},{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/topics\\\/#listItem\",\"position\":2,\"name\":\"Topics\",\"item\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/topics\\\/\",\"nextItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/forums\\\/topic\\\/lumerical-cluster-job-crashes-arbitrarily\\\/#listItem\",\"name\":\"Lumerical cluster job crashes arbitrarily\"},\"previousItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum#listItem\",\"name\":\"Home\"}},{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/forums\\\/topic\\\/lumerical-cluster-job-crashes-arbitrarily\\\/#listItem\",\"position\":3,\"name\":\"Lumerical cluster job crashes arbitrarily\",\"previousItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/topics\\\/#listItem\",\"name\":\"Topics\"}}]},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/#organization\",\"name\":\"Ansys Learning Forum\",\"description\":\"Ansys Innovation Space\",\"url\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/forums\\\/topic\\\/lumerical-cluster-job-crashes-arbitrarily\\\/#webpage\",\"url\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/forums\\\/topic\\\/lumerical-cluster-job-crashes-arbitrarily\\\/\",\"name\":\"Lumerical cluster job crashes arbitrarily | Ansys Learning Forum\",\"description\":\"I run many parallel jobs on a cluster. Sometimes - for no apparent reason, the job crashes citing \\\"std::bad_alloc\\\", but no more information. The RAM is sufficient (I have verified this, I have allocated 10x the RAM that the requirements ask for - and have monitored that this is not reached via observing using htop).\",\"inLanguage\":\"en-US\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/#website\"},\"breadcrumb\":{\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/forums\\\/topic\\\/lumerical-cluster-job-crashes-arbitrarily\\\/#breadcrumblist\"},\"datePublished\":\"2024-04-15T22:44:24+00:00\",\"dateModified\":\"2024-04-15T22:44:24+00:00\"},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/#website\",\"url\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/\",\"name\":\"Ansys Learning Forum\",\"description\":\"Ansys Innovation Space\",\"inLanguage\":\"en-US\",\"publisher\":{\"@id\":\"https:\\\/\\\/innovationspace.ansys.com\\\/forum\\\/#organization\"}}]}\n\t\t<\/script>\n\t\t<!-- All in One SEO -->\n\n","aioseo_head_json":{"title":"Lumerical cluster job crashes arbitrarily | Ansys Learning Forum","description":"I run many parallel jobs on a cluster. Sometimes - for no apparent reason, the job crashes citing \"std::bad_alloc\", but no more information. The RAM is sufficient (I have verified this, I have allocated 10x the RAM that the requirements ask for - and have monitored that this is not reached via observing using htop).","canonical_url":"https:\/\/innovationspace.ansys.com\/forum\/forums\/topic\/lumerical-cluster-job-crashes-arbitrarily\/","robots":"max-image-preview:large","keywords":"","webmasterTools":{"miscellaneous":""},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"BreadcrumbList","@id":"https:\/\/innovationspace.ansys.com\/forum\/forums\/topic\/lumerical-cluster-job-crashes-arbitrarily\/#breadcrumblist","itemListElement":[{"@type":"ListItem","@id":"https:\/\/innovationspace.ansys.com\/forum#listItem","position":1,"name":"Home","item":"https:\/\/innovationspace.ansys.com\/forum","nextItem":{"@type":"ListItem","@id":"https:\/\/innovationspace.ansys.com\/forum\/topics\/#listItem","name":"Topics"}},{"@type":"ListItem","@id":"https:\/\/innovationspace.ansys.com\/forum\/topics\/#listItem","position":2,"name":"Topics","item":"https:\/\/innovationspace.ansys.com\/forum\/topics\/","nextItem":{"@type":"ListItem","@id":"https:\/\/innovationspace.ansys.com\/forum\/forums\/topic\/lumerical-cluster-job-crashes-arbitrarily\/#listItem","name":"Lumerical cluster job crashes arbitrarily"},"previousItem":{"@type":"ListItem","@id":"https:\/\/innovationspace.ansys.com\/forum#listItem","name":"Home"}},{"@type":"ListItem","@id":"https:\/\/innovationspace.ansys.com\/forum\/forums\/topic\/lumerical-cluster-job-crashes-arbitrarily\/#listItem","position":3,"name":"Lumerical cluster job crashes arbitrarily","previousItem":{"@type":"ListItem","@id":"https:\/\/innovationspace.ansys.com\/forum\/topics\/#listItem","name":"Topics"}}]},{"@type":"Organization","@id":"https:\/\/innovationspace.ansys.com\/forum\/#organization","name":"Ansys Learning Forum","description":"Ansys Innovation Space","url":"https:\/\/innovationspace.ansys.com\/forum\/"},{"@type":"WebPage","@id":"https:\/\/innovationspace.ansys.com\/forum\/forums\/topic\/lumerical-cluster-job-crashes-arbitrarily\/#webpage","url":"https:\/\/innovationspace.ansys.com\/forum\/forums\/topic\/lumerical-cluster-job-crashes-arbitrarily\/","name":"Lumerical cluster job crashes arbitrarily | Ansys Learning Forum","description":"I run many parallel jobs on a cluster. Sometimes - for no apparent reason, the job crashes citing \"std::bad_alloc\", but no more information. The RAM is sufficient (I have verified this, I have allocated 10x the RAM that the requirements ask for - and have monitored that this is not reached via observing using htop).","inLanguage":"en-US","isPartOf":{"@id":"https:\/\/innovationspace.ansys.com\/forum\/#website"},"breadcrumb":{"@id":"https:\/\/innovationspace.ansys.com\/forum\/forums\/topic\/lumerical-cluster-job-crashes-arbitrarily\/#breadcrumblist"},"datePublished":"2024-04-15T22:44:24+00:00","dateModified":"2024-04-15T22:44:24+00:00"},{"@type":"WebSite","@id":"https:\/\/innovationspace.ansys.com\/forum\/#website","url":"https:\/\/innovationspace.ansys.com\/forum\/","name":"Ansys Learning Forum","description":"Ansys Innovation Space","inLanguage":"en-US","publisher":{"@id":"https:\/\/innovationspace.ansys.com\/forum\/#organization"}}]},"og:locale":"en_US","og:site_name":"Ansys Learning Forum | Ansys Innovation Space","og:type":"article","og:title":"Lumerical cluster job crashes arbitrarily | Ansys Learning Forum","og:description":"I run many parallel jobs on a cluster. Sometimes - for no apparent reason, the job crashes citing &quot;std::bad_alloc&quot;, but no more information. The RAM is sufficient (I have verified this, I have allocated 10x the RAM that the requirements ask for - and have monitored that this is not reached via observing using htop).","og:url":"https:\/\/innovationspace.ansys.com\/forum\/forums\/topic\/lumerical-cluster-job-crashes-arbitrarily\/","article:published_time":"2024-04-15T22:44:24+00:00","article:modified_time":"2024-04-15T22:44:24+00:00","twitter:card":"summary_large_image","twitter:title":"Lumerical cluster job crashes arbitrarily | Ansys Learning Forum","twitter:description":"I run many parallel jobs on a cluster. Sometimes - for no apparent reason, the job crashes citing &quot;std::bad_alloc&quot;, but no more information. The RAM is sufficient (I have verified this, I have allocated 10x the RAM that the requirements ask for - and have monitored that this is not reached via observing using htop)."},"aioseo_meta_data":{"post_id":"362344","title":null,"description":null,"keywords":null,"keyphrases":null,"primary_term":null,"canonical_url":null,"og_title":null,"og_description":null,"og_object_type":"default","og_image_type":"content","og_image_url":null,"og_image_width":null,"og_image_height":null,"og_image_custom_url":null,"og_image_custom_fields":null,"og_video":null,"og_custom_url":null,"og_article_section":null,"og_article_tags":null,"twitter_use_og":true,"twitter_card":"default","twitter_image_type":"default","twitter_image_url":null,"twitter_image_custom_url":null,"twitter_image_custom_fields":null,"twitter_title":null,"twitter_description":null,"schema":{"blockGraphs":[],"customGraphs":[],"default":{"data":{"Article":[],"Course":[],"Dataset":[],"FAQPage":[],"Movie":[],"Person":[],"Product":[],"ProductReview":[],"Car":[],"Recipe":[],"Service":[],"SoftwareApplication":[],"WebPage":[]},"graphName":"","isEnabled":true},"graphs":[]},"schema_type":"default","schema_type_options":null,"pillar_content":false,"robots_default":true,"robots_noindex":false,"robots_noarchive":false,"robots_nosnippet":false,"robots_nofollow":false,"robots_noimageindex":false,"robots_noodp":false,"robots_notranslate":false,"robots_max_snippet":null,"robots_max_videopreview":null,"robots_max_imagepreview":"large","priority":null,"frequency":null,"local_seo":null,"limit_modified_date":false,"created":"2024-10-28 16:47:11","updated":"2024-10-28 16:47:11","ai":null,"breadcrumb_settings":null,"seo_analyzer_scan_date":null},"aioseo_breadcrumb":"<div class=\"aioseo-breadcrumbs\"><span class=\"aioseo-breadcrumb\">\n\t\t\t<a href=\"https:\/\/innovationspace.ansys.com\/forum\" title=\"Home\">Home<\/a>\n\t\t<\/span><span class=\"aioseo-breadcrumb-separator\">&raquo;<\/span><span class=\"aioseo-breadcrumb\">\n\t\t\t<a href=\"https:\/\/innovationspace.ansys.com\/forum\/topics\/\" title=\"Topics\">Topics<\/a>\n\t\t<\/span><span class=\"aioseo-breadcrumb-separator\">&raquo;<\/span><span class=\"aioseo-breadcrumb\">\n\t\t\tLumerical cluster job crashes arbitrarily\n\t\t<\/span><\/div>","aioseo_breadcrumb_json":[{"label":"Home","link":"https:\/\/innovationspace.ansys.com\/forum"},{"label":"Topics","link":"https:\/\/innovationspace.ansys.com\/forum\/topics\/"},{"label":"Lumerical cluster job crashes arbitrarily","link":"https:\/\/innovationspace.ansys.com\/forum\/forums\/topic\/lumerical-cluster-job-crashes-arbitrarily\/"}],"acf":[],"custom_fields":[{"0":{"_bbp_subscription":["343740","4274"],"_bbp_author_ip":["23.206.193.146"]," _bbp_last_reply_id":["0"]," _bbp_likes_count":["0"],"_btv_view_count":["568"],"_bbp_topic_status":["unanswered"],"_bbp_topic_id":["362344"],"_bbp_forum_id":["27833"],"_bbp_engagement":["4274","343740"],"_bbp_voice_count":["2"],"_bbp_reply_count":["6"],"_bbp_last_reply_id":["363618"],"_bbp_last_active_id":["363618"],"_bbp_last_active_time":["2024-04-22 21:06:03"]},"test":"vn95cornell-edu"}],"_links":{"self":[{"href":"https:\/\/innovationspace.ansys.com\/forum\/wp-json\/wp\/v2\/topics\/362344","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/innovationspace.ansys.com\/forum\/wp-json\/wp\/v2\/topics"}],"about":[{"href":"https:\/\/innovationspace.ansys.com\/forum\/wp-json\/wp\/v2\/types\/topic"}],"version-history":[{"count":0,"href":"https:\/\/innovationspace.ansys.com\/forum\/wp-json\/wp\/v2\/topics\/362344\/revisions"}],"wp:attachment":[{"href":"https:\/\/innovationspace.ansys.com\/forum\/wp-json\/wp\/v2\/media?parent=362344"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}