# Segfaults in arm64 environment

**URL:** <https://travis-ci.community/t/segfaults-in-arm64-environment/5617>\
**Category:** Multi CPU Architecture\
**Tags:** build-env, bug\
**Created:** [October 22, 2019, 11:59am UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617 "2019-10-22T11:59:47Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![noloader](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/noloader/32/2971_2.png) [@noloader](https://travis-ci.community/u/noloader)\
**Post date:** [October 22, 2019, 11:59am UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/1 "2019-10-22T11:59:47Z")

</div>

Hi Everyone,

We successfully cut-in arm64 testing with GCC and Clang late last week. The Linux testing uses Xenial images. All CI testing passed.

This week we notice there are unexplained segfaults in arm64. An example is [here](https://travis-ci.org/noloader/cryptopp/builds/601210611). A typical message is shown below.

```auto
Testing SymmetricCipher algorithm RabbitWithIV.
............................................................
............................................................
.........../home/travis/.travis/functions: line 134: 2018 S
egmentation fault

```

The segfault moves around. Sometimes one set of tests fail, and at other times another set of tests fail.

We reverted a few commits to go back to the last known good but the arm64 segfaults persist. Our last known good is [here](https://travis-ci.org/noloader/cryptopp/builds/600672339).

We can’t duplicate the arm64 segfaults at the [compile farm](https://cfarm.tetaneutral.net/machines/list/) (GCC117 and GCC118), and we can’t duplicate it on four aarch64 dev-boards. We also cannot duplicate it on other arch’es and environments, like x86, x86\_64, arm, ppc64le or ppc64be.

I’m beginning to suspect the script `/home/travis/.travis/functions` or something similar.

From the build information at the head of the output, I’ve gotten this far:

- lask known good: `travis-build version: 2f1f818b6`
- segfaults: `travis-build version: a91ac50bd`

One of the side effects of the build version change:

- lask known good: `gcc (Ubuntu/Linaro 7.4.0-1ubuntu1~18.04.1) 7.4.0`
- segfaults: `gcc (Ubuntu/Linaro 5.4.0-6ubuntu1~16.04.11) 5.4.0`

My first question is, what changed on Sunday or Monday in the arm64 environment?

My second question is, how do we work around the changes?

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![Michal](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/michal/32/2848_2.png) [@Michal](https://travis-ci.community/u/Michal)\
**Post date:** [October 23, 2019, 10:31am UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/2 "2019-10-23T10:31:46Z")

</div>

Hi @noloader

Thanks for detailed feedback and happy to see you using Arm builds!

One request: are you able to verify if the segfaults occur with `dist: bionic` ?

Before Monday all `dist` references ended up with Ubuntu Bionic OS image anyway - the OS reported in build job log Build Environment section is taken from .travis.yml rather than from actual image.

On Monday we’ve added actual Xenial OS image in order to allow builds run within LXD container on proper target OS. So your `dist: xenial` started to be actually built on Xenial run as an LXD container on LXD host (the LXD host is Bionic at the moment ).  
(all of above in context of Arm64 builds of course)

So - by checking if segfaults occur on `dist: bionic` along with info you provided so far could help narrow down the cause.

Best Regards  
Michał

---

<div class="post-metadata">

**Author:** ![Michal](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/michal/32/2848_2.png) [@Michal](https://travis-ci.community/u/Michal)\
**Post date:** [November 20, 2019, 12:02pm UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/3 "2019-11-20T12:02:55Z")

</div>

Hi @noloader!

Is this segmentation fault on arm64 still occurring for you? I’ve seen you’ve been adopting also IBM power and Z targets. The reason I’m asking is, that at the moment our [cache](https://docs.travis-ci.com/user/caching/) meant to keep items between builds and beta [workspaces](https://docs.travis-ci.com/user/using-workspaces/) (artifacts between jobs in a build) should work correctly also for different architectures. If you have a binary, for which the segmentation fault is reproducible, it’d help us to debug it. Maybe such binary could be deployed using recent [DPL2](https://docs.travis-ci.com/user/deployment-v2) to a place, from where we could download it from?

---

<div class="post-metadata">

**Author:** ![Marco](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/marco/32/3441_2.png) [@Marco](https://travis-ci.community/u/Marco)\
**Post date:** [November 20, 2019, 3:57pm UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/4 "2019-11-20T15:57:56Z")

</div>

Hi, we’ve also been seeing segfaults on arm64. I can’t reproduce them locally, but we suspect that the issue might be overheating of the hardware or a hardware fault. See [https://github.com/bitcoin/bitcoin/issues/17481](https://github.com/bitcoin/bitcoin/issues/17481)

Compiling will use 100% CPU, so if the heat is not properly dealt with, it could lead to intermittent hardware issues.

---

<div class="post-metadata">

**Author:** ![Marco](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/marco/32/3441_2.png) [@Marco](https://travis-ci.community/u/Marco)\
**Post date:** [November 21, 2019, 8:31pm UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/5 "2019-11-21T20:31:42Z")

</div>

Here is another intermittent segfault in the compiler: [https://travis-ci.org/MarcoFalke/bitcoin-core/jobs/615121672#L8342](https://travis-ci.org/MarcoFalke/bitcoin-core/jobs/615121672#L8342)

Edit: And another one: [https://travis-ci.org/bitcoin/bitcoin/jobs/615102790#L13961](https://travis-ci.org/bitcoin/bitcoin/jobs/615102790#L13961)

---

<div class="post-metadata">

**Author:** ![noloader](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/noloader/32/2971_2.png) [@noloader](https://travis-ci.community/u/noloader)\
**Post date:** [November 22, 2019, 12:22am UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/6 "2019-11-22T00:22:11Z")

</div>

We think we narrowed it down to Xenial images and/or the compiler. bitcoin-core seems to have the same problem. From Build System Information around line 7:

```auto
Build language: minimal
Build group: stable
Build dist: xenial
Build id: 615102787
Job id: 615102790
Runtime kernel version: 5.3.0-22-generic
travis-build version: a09969ae2

```

You can probably sidestep the problem by switching to Bionic images. Just use `dist: bionic` in you Travis yml file.

I suspect a Xenial dist-upgrade will also fix the issue, but I don’t know for sure. I seem to recall the Travis docs ask folks to avoid dist-upgrade, so I’m not even sure you can dist-upgrade a Travis image.

---

<div class="post-metadata">

**Author:** ![Marco](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/marco/32/3441_2.png) [@Marco](https://travis-ci.community/u/Marco)\
**Post date:** [November 22, 2019, 3:04pm UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/7 "2019-11-22T15:04:11Z")

</div>

We use the xenial travis image, but run everything in a bionic docker. See [https://travis-ci.org/MarcoFalke/bitcoin-core/jobs/615121672#L116](https://travis-ci.org/MarcoFalke/bitcoin-core/jobs/615121672#L116)

Because it is only intermittent, I think the issue is with the hardware/overheating as mentioned previously.

---

<div class="post-metadata">

**Author:** ![yzyuestc](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/yzyuestc/32/4694_2.png) [@yzyuestc](https://travis-ci.community/u/yzyuestc)\
**Post date:** [December 16, 2019, 1:38am UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/8 "2019-12-16T01:38:05Z")

</div>

Hi, I met similar issue on travis aarch64 with ubuntu bionic. Please see the build report:  
[https://travis-ci.org/ovsrobot/ovs/jobs/621434216#L1999](https://travis-ci.org/ovsrobot/ovs/jobs/621434216#L1999)

I saw some similar issues in the community. Does any know the cause and any solution to avoid it?

---

<div class="post-metadata">

**Author:** ![yzyuestc](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/yzyuestc/32/4694_2.png) [@yzyuestc](https://travis-ci.community/u/yzyuestc)\
**Post date:** [January 2, 2020, 8:23am UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/9 "2020-01-02T08:23:39Z")

</div>

Hi,

As the gcc segment faults issue occurred unexpected, I would like to do some investigation for the possible cause. I checked the report at [https://travis-ci.org/ovsrobot/ovs/jobs/621434216#L2005](https://travis-ci.org/ovsrobot/ovs/jobs/621434216#L2005). I found a log: “See \<file:///usr/share/doc/gcc-5/README.Bugs\> for instructions.”

Is there a way to fetch the bug report at /usr/share/doc/gcc-5/README.Bugs in the previous build job? Any response is appreciated.

---

<div class="post-metadata">

**Author:** ![Damian](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/damian/32/3625_2.png) [@Damian](https://travis-ci.community/u/Damian)\
**Post date:** [January 2, 2020, 8:52am UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/10 "2020-01-02T08:52:43Z")

</div>

Hi @yzyuestc  
Sorry I am afraid there is no way, container is destroyed after job is done.

---

<div class="post-metadata">

**Author:** ![native-api](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/native-api/32/430_2.png) [@native-api](https://travis-ci.community/u/native-api)\
**Post date:** [January 2, 2020, 10:56am UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/11 "2020-01-02T10:56:48Z")

</div>

> [@yzyuestc](#):
>
> Is there a way to fetch the bug report at /usr/share/doc/gcc-5/README.Bugs in the previous build job?

Since this is a static file, thus should be the same for every job run in the same environment, you can either `cat` it (or upload somewhere) in another build, or get the corresponding package online and extract it from there.  
(And, just to be clear, this is not a bug report but a part of the `gcc` package’s documentation – perhaps with instructions how to file bug reports.)

---

<div class="post-metadata">

**Author:** ![Michal](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/michal/32/2848_2.png) [@Michal](https://travis-ci.community/u/Michal)\
**Post date:** [January 8, 2020, 5:40pm UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/12 "2020-01-08T17:40:10Z")

</div>

@yzyuestc

Did you manage to capture anything more?

@Marco @noloader@yzyuestc - thank yopu for your reports and effort so far.  
This is kind of vanishing point for us.

1. Happens occasionally/on specific builds, but clearly often enough to be a stability issue while it’s not happening on different hardware/outside of LXD container
2. Happens with `dist: xenial` so far and gcc7 , however there is at least one confirmed case of problem occurring on `bionic` (see OvS)

The hardware/temperature issue - we cannot verify it on our end easily. We’re not owning the infrastructure. I’ll ask around though.

The gcc version - can you check if the issue re-occurs with most up-to date gcc version installed at the beginning of the job?

The xenial vs bionic - LXD container is running a Xenial Ubuntu, however the kernel itself is shared by host (it’s the core concept for this container), which is 5.x from Bionic host. As far as we saw and know, this combination has been stable on arm64, yet seems to be a trigger for the initial case. Verification requires a binary being result/partial result of failed job. The binary - is this possible for any of you to provide/test locally specifically a binary parts uploaded/deployed from failed build? Does it fail at your local environments?

---

<div class="post-metadata">

**Author:** ![Michal](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/michal/32/2848_2.png) [@Michal](https://travis-ci.community/u/Michal)\
**Post date:** [January 9, 2020, 1:21pm UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/13 "2020-01-09T13:21:29Z")

</div>

@noloader @Marco @yzyuestc

We have updated setup on our end - to which machines the arm jobs are redirected - to indirectly verify one suspicious instance. The change is effective as of 2020 Jan 9th, 13:00 UTCZ. Could you please let us know, do you still observe random segmentation faults after that time?

---

<div class="post-metadata">

**Author:** ![yzyuestc](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/yzyuestc/32/4694_2.png) [@yzyuestc](https://travis-ci.community/u/yzyuestc)\
**Post date:** [January 10, 2020, 1:34am UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/14 "2020-01-10T01:34:51Z")

</div>

Hi Michal, @Michal

Thanks for your effort and update. We will try some new builds to observer whether the segfault issue will occur or not.

Thank you again!

---

<div class="post-metadata">

**Author:** ![noloader](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/noloader/32/2971_2.png) [@noloader](https://travis-ci.community/u/noloader)\
**Post date:** [January 10, 2020, 2:26am UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/15 "2020-01-10T02:26:21Z")

</div>

Hi Michal,

We have not experienced the issue since switching to Bionic images.

Jeff

---

<div class="post-metadata">

**Author:** ![Michal](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/michal/32/2848_2.png) [@Michal](https://travis-ci.community/u/Michal)\
**Post date:** [January 10, 2020, 1:41pm UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/16 "2020-01-10T13:41:18Z")

</div>

Hi Jeff  
Appreciated, thank you!

Hi @yzyuestc  
We will wait for your feedback before continuing with potential changes. We’ve been in touch with infrastructure provider to solve it permanently and would like to be sure if now everything works stable.

---

<div class="post-metadata">

**Author:** ![yzyuestc](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/yzyuestc/32/4694_2.png) [@yzyuestc](https://travis-ci.community/u/yzyuestc)\
**Post date:** [January 16, 2020, 3:13am UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/17 "2020-01-16T03:13:24Z")

</div>

Hi Michal,

The segment fault issue has not been reproduced since the time your replied. It seems everything works fine on my side.

---

<div class="post-metadata">

**Author:** ![Michal](https://sea1.discourse-cdn.com/flex015/user_avatar/travis-ci.community/michal/32/2848_2.png) [@Michal](https://travis-ci.community/u/Michal)\
**Post date:** [January 16, 2020, 12:27pm UTC](https://travis-ci.community/t/segfaults-in-arm64-environment/5617/18 "2020-01-16T12:27:04Z")

</div>

Thank you @yzyuestc.  
One down. We will cure the patient 😉

Now to the other on segfaults 😉
